Higher Education Malpractice: curving grades

If there is one thing that university faculty and administrators could do today to demonstrate their commitment to inclusion, not to mention teaching and learning over sorting and status, it would be to ban curve-based, norm-referenced grading. Many obstacles exist to the effective inclusion and success of students from underrepresented (and underserved) groups in science and related programs.  Students and faculty often, and often correctly, perceive large introductory classes as “weed out” courses preferentially impacting underrepresented students. In the life sciences, many of these courses are “out-of-major” requirements, in which students find themselves taught with relatively little regard to the course’s relevance to bio-medical careers and interests. Often such out-of-major requirements spring not from a thoughtful decision by faculty as to their necessity, but because they are prerequisites for post-graduation admission to medical or graduate school. “In-major” instructors may not even explicitly incorporate or depend upon the materials taught in these out-0f-major courses – rare is the undergraduate molecular biology degree program that actually calls on students to use calculus or a working knowledge of physics, despite the fact that such skills may be relevant in certain biological contexts – see Magnetofiction – A Reader’s Guide.  At the same time, those teaching “out of major” courses may overlook the fact that many (and sometimes most) of their students are non-chemistry, non-physics, and/or non-math majors.  The result is that those teaching such classes fail to offer a doorway into the subject matter to any but those already comfortable with it. But reconsidering the design and relevance of these courses is no simple matter.  Banning grading on a curve, on the other  hand, can be implemented overnight (and by fiat if necessary). 

 So why ban grading on a curve?  First and foremost, it would put faculty and institutions on record as valuing student learning outcomes (perhaps the best measure of effective teaching) over the sorting of students into easy-to-judge groups.  Second, there simply is no pedagogical justification for curved grading, with the possible exception of providing a kludgy fix to correct for poorly designed examinations and courses. There are more than enough opportunities to sort students based on their motivation, talent, ambition, “grit,” and through the opportunities they seek after and successfully embraced (e.g., through volunteerism, internships, and independent study projects). 

The negative impact of curving can be seen in a recent paper by Harris et al,  (Reducing achievement gaps in undergraduate general chemistry …), who report a significant difference in overall student inclusion and subsequent success based on a small grade difference between a C, which allows a student to proceed with their studies (generally as successfully as those with higher grades) and a C-minus, which requires them to retake the course before proceeding (often driving them out of the major).  Because Harris et al., analyzed curved courses, a subset of students cannot escape these effects.  And poor grades disproportionately impact underrepresented and underserved groups – they say explicitly “you do not belong” rather than “how can I help you learn”.   

Often naysayers disparage efforts to improve course design as “dumbing down” the course, rather than improving it.  In many ways this is a situation analogous to blaming patients for getting sick or not responding to treatment, rather than conducting an objective analysis of the efficacy of the treatment.  If medical practitioners had maintained this attitude, we would still be bleeding patients and accepting that more than a third are fated to die, rather than seeking effective treatments tailored to patients’ actual diseases – the basis of evidence-based medicine.  We would have failed to develop antibiotics and vaccines – indeed, we would never have sought them out. Curving grades implies that course design and delivery are already optimal, and the fate of students is predetermined because only a percentage can possibly learn the material.  It is, in an important sense, complacent quackery.

Banning grading on a curve, and labelling it for what it is – educational malpractice – would also change the dynamics of the classroom and might even foster an appreciation that a good teacher is one with the highest percentage of successful students, e.g. those who are retained in a degree program and graduate in a timely manner (hopefully within four years). Of course, such an alternative evaluation of teaching would reflect a department’s commitment to construct and deliver the most engaging, relevant, and effective educational program. Institutional resources might even be used to help departments generate more objective, instructor-independent evaluations of learning outcomes, in part to replace the current practice of student-based opinion surveys, which are often little more than measures of popularity.  We might even see a revolution in which departments compete with one another to maximize student inclusion, retention, and outcomes (perhaps even to the extent of applying pressure on the design and delivery of “out of major” required courses offered by other departments).  

“All a pipe dream” you might say, but the available data demonstrates that resources spent on rethinking course design, including engagement and relevance, can have significant effects on grades, retention, time to degree, and graduation rates.  At the risk of being labeled as self-promoting, I offer the following to illustrate the possibilities: working with Melanie Cooper at Michigan State University, we have built such courses in general and organic chemistry and documented their impact, see Evaluating the extent of a large-scale transformation in gateway science courses.

Perhaps we should be encouraging students to seek out legal representation to hold institutions (and instructors) accountable for detrimental practices, such as grading on a curve.  There might even come a time when professors and departments would find it prudent to purchase malpractice insurance if they insist on retaining and charging students for ineffective educational strategies.(1)  

Acknowledgements: Thanks to daughter Rebecca who provided edits and legal references and Melanie Cooper who inspired the idea. Educate! image from the Dorian De Long Arts & Music Scholarship site.

(1) One cannot help but wonder if such conduct could ever rise to the level of fraud. See, e.g., Bristol Bay Productions, LLC vs. Lampack, 312 P.3d 1155, 1160 (Colo. 2013) (“We have typically stated that a plaintiff seeking to prevail on a fraud claim must establish five elements: (1) that the defendant made a false representation of a material fact; (2) that the one making the representation knew it was false; (3) that the person to whom the representation was made was ignorant of the falsity; (4) that the representation was made with the intention that it be acted upon; and (5) that the reliance resulted in damage to the plaintiff.”).

Thinking about biological thinking: Steady state, half-life & response dynamics

Insights into student thinking & course design, part of the biofundamentals project. 

Something that often eludes both instructors and instructional researchers is a clear appreciation of what it is that students do and do not know, what ideas they can and cannot call upon to solve problems and generate clear, coherent, and plausible explanations. What information – thought to have been presented effectively through past instruction, appears to be unavailable to students. As an example, few instructors would believe that students completing college level chemistry could possibly be confused about the differences between covalent and non-covalent molecular interactions, yet there is good evidence that they are (Williams et al., 2015). Unless these ideas, together with their  conceptual bases and practical applications, are explicitly called out in the design and implementation of instructional materials, they often fail to become a working (relevant) part of the students’ conceptual tool-kit.   

To identify ideas involved in understanding biological systems, we are using an upper division undergraduate course in developmental biology (blog link) to provide context; this is a final “capstone” junior/senior level course that comes after students have completed multiple required courses in chemistry and biology.  Embryonic development integrates a range of molecular level processes, including the control of gene expression, cellular morphology and dynamics, through intrinsic and extrinsic signaling systems.   

A key aspect of the course’s design is the use of formative assessment activities delivered through the beSocratic system. These activities generally include parts in which students are asked to draw a graph or diagram. Students are required to complete tasks before the start of each class meeting; their responses are used to inform in-class discussions, a situation akin to reviewing game film and coaching in sports. Analysis of student drawings and comments, carried out in collaboration with Melanie Cooper and her group at Michigan State University, can reveal unexpected aspects of students’ thinking (e.g. Williams et al., 2015). What emerges from this Socratic give and take is an improved appreciation of the qualities of the tasks that engage students (as well as those that do not), and insights into how students analyze specific tasks, what sets of ideas they see as necessary and which necessary ideas they ignore when generating explanatory and predictive models. Most importantly, they can reveal flaws in how necessary ideas are developed. While at an admittedly early stage in the project, here I sketch out some preliminary findings: the first of these deal with steady state concentration and response dynamics.

The ideas of steady state concentration and pathway dynamics were identified by Loertscher et al (2014)as two of five “threshold concepts” in  biochemistry and presumably molecular biology as well. Given the non-equilibrium nature of biological systems, we consider the concentration of a particular molecule in a cell in dynamic terms, a function of its rate of synthesis (or importation from the environment) together with its rate of breakdown.  On top of this dynamic, the activity of existing molecules can be regulated through various post-translational mechanisms.  All of the populations of molecules within a cell or organism have a characteristic steady state concentration with the exception of genomic DNA, which while synthesized is not, in living organisms, degraded, although it is repaired.

In biological systems, molecules are often characterized by their “half life” but this can be confusing, since it is quite different from the way the term is used in physics, where students are likely to first be introduced to it.[1]  Echos from physics can imply that a molecule’s half-life is an intrinsic feature of the molecule, rather than of the system in which the molecule finds itself.  The equivalent of half-life would be doubling time, but these terms make sense only under specific conditions.  In a system in which synthesis has stopped (synthesis rate = 0) the half life is the time it takes for the number of molecules in the system to decrease by 50%, while in the absence of degradation (degradation rate = 0), the doubling time is the time it takes to double the number of molecules in the system.  Both degradation and synthesis rates are regulateable and can vary, often dramatically, in response to various stimuli.

In the case of RNA and polypeptide levels, the synthesis rate is determined by many distinct processes, including effective transcription factor concentrations, the signals that activate transcription factors, rates of binding of transcription factors to transcription factor binding sites (which can involve both DNA sequences and other proteins), as well as relevant binding affinities, and the rates associated with the recruitment and activation of DNA-dependent, RNA polymerase. Once activated, the rate of gene specific RNA synthesis will be influenced  by the rate of RNA polymerization (nucleotide bases added per second) and the length of the RNA molecules synthesized.  In eukaryotes, the newly formed RNA will generally need to have introns removed through interactions with splicing machinery, as well as other  post-transcriptional reactions, after which the processed RNA will be transported from the nucleus to the cytoplasm through the nuclear pore complex. In the cytoplasm there are rates associated with the productive interaction of RNAs with the translational machinery (ribosomes and associated factors), and the rate at which polypeptide synthesis occurs (amino acids added per second) together with the length of the polypeptide synthesized (given that things are complicated enough, I will ignore processes such as those associated with the targeting of membrane proteins and codon usage, although these will be included in a new chapter in biofundamentals reasonably soon, I hope). On the degradative side, there are rates associated with interactions with nucleases (that breakdown RNAs) and proteinases (that breakdown polypeptides).  These processes are energy requiring; generally driven by reactions coupled to the hydrolysis of adenosine triphosphate (ATP). 

That these processes matter is illustrated nicely in work from Harima and colleagues (2014).   The system, involved in the segmentation of the anterior region of the presomitic mesoderm, responds to signaling by activating the Hes7 gene, while the Hes7 gene product act to inhibit Hes7 gene expression. The result is an oscillatory response that is “tuned” by the length of the transcribed region (RNA length). This can be demonstrated experimentally by generating mice in which two of the genes three introns (Hes7-3) or all three introns (intron-less) are removed. Removing introns changes the oscillatory behavior of the system (Hes7 mRNA -blue and Hes7 protein – green)(Harima et al., 2013).

In the context of developmental biology, we use beSocratic activities to ask students to consider a molecule’s steady state concentration as a function of its synthesis and degradation rates, and to predict how the system would change when one or the other is altered. These ideas were presented in the context of observations by Schwanhausser et al (2011) that large discrepancies between steady state RNA and polypeptide concentrations are common and that there is an absence of a correlation between RNA and polypeptide half-lives (we also use these activities to introduce the general idea of correlation). In their responses, it was common to see students’ linking high steady state concentrations exclusively to long half-lives. Ask to consider the implications in terms of system responsiveness (in the specific context of a positively-acting transcription factor and target gene expression), students often presumed that a longer half-life would lead to higher steady state concentration which in turn would lead to increased target gene expression, primarily because collisions between the transcription factor and its DNA-binding sites would increase, leading to higher levels of target gene expression. This is an example of a p-prim (Hammer, 1996) – the heuristic that “more is more”, a presumption that is applicable to many systems. 

In biological systems, however, this is generally not the case – responses “saturate”, that is  increasing transcription factor concentration (or activity) above a certain level generally does not lead to a proportionate, or any increase in target gene expression. We would not call this a misconception, because this is an example of an idea that is useful in many situations, but generally isn’t in biological systems – where responses are generally inherently limited. The ubiquity and underlying mechanisms of response saturation need to be presented explicitly, and its impact on various processes reinforced repeatedly, preferably by having students use them to solve problems or construct plausible explanations. A related phenomenon that students seemed not to recognize involves the non-linearity of the initial response to a stimulus, in this case, the concentration of transcription factor below which target gene expression is not observed (or it may occur, but only transiently or within a few cells in the population, so as to be undetectable by the techniques used).

So what ideas do students need to call upon when they consider steady state concentration, how it changes, and the impact of such changes on system behavior?  It seems we need to go beyond synthesis and degradation rates and include the molecular processes associated with setting the system’s response onset and saturation concentrations.  First we need to help students appreciate why such behaviors (onset and saturation) occur – why doesn’t target gene expression begin as soon as a transcription factor appears in a cell?  Why does gene expression level off when transcription factor concentrations rise above a certain level?  The same questions apply to the types of threshold behaviors often associated with signaling systems.  For example, in quorum sensing among unicellular organisms, the response of cells to the signal occurs over a limited concentration range, from off to full on.  A related issue is associated with morphogen gradients (concentration gradients over space rather than time), in which there are multiple distinct types of “threshold” responses. One approach might be to develop a model in which we set the onset concentration close to the saturation concentration. The difficulty (or rather instructional challenge) here is that these are often complex processes involving cooperative as well as feedback interactions.

Our initial approach to steady state and thresholds has been to build activities based on the analysis of a regulatory network presented by Saka and Smith (2007), an analysis based on studies of early embryonic development in the frog Xenopus laevis. We chose the system because of its simplicity, involving only four components (although there are many other proteins associated with the actual system).  Saka and Smith modeled the regulatory network controlling the expression of the transcription factor proteins Goosecoid (Gsc) and Brachyury (Xbra) in response to the secreted signaling protein activin (↓), a member of

the TGFβ superfamily of secreted signaling proteins (see Li and Elowitz, 2019).   The network involves the positive action of Xbra on the gene encoding the transcription factor protein Xom.  The system’s behavior depends on the values of various parameters, parameters that include response to activator (Activin), rates of synthesis and the half-lives of Gsc, Xbra, and Xom, and the degrees of regulatory cooperativity and responsiveness.

Depending upon these parameters, the system can produce a range of complex responses.  In different regimes (→),  increasing concentrations of activin (M) can lead, initially, to increasing, but mutually exclusive, expression of either Xba (B) or Gsc (A) as well as sharp transitions in which expression flips from one to the other, as Activin concentration increases, after which the response saturates. There are also conditions at very low Activin concentration (marked by ↑) in which both Xbra and Gsc are expressed at low levels, a situation that students are asked to explain.

Lessons learned: Based on their responses, captured through beSocratic and revealed during in class discussions, it appears that there is a need to be more explicit (early in the course, and perhaps the curriculum as well) when considering the mechanisms associated with response onset and saturation, in the context of how changes in the concentrations of regulatory factors (through changes in synthesis, turn-over, and activity) impact system responses. This may require a more quantitative approach to molecular dynamics and system behaviors. Here we may run into a problem, the often phobic responses of biology majors (and many faculty) to mathematical analyses.  Even the simplest of models, such as that of Saka and Smith, require a consideration of factors generally unfamiliar to students, concepts and skills that may well not be emphasized or mastered in prerequisite courses. The trick is to define realistic, attainable, and non-trivial goals – we are certainly not going to succeed in getting late stage molecular biology students with rudimentary math skills to solve systems of differential equations in a developmental biology course.  But perhaps we can build up the instincts needed to appreciate the molecular processes involved in the behavior of systems whose behavior evolves overtime in response to various external signals (which is, of course, pretty much every biological system).


[1] A similar situation exists in the context of the term “spontaneous” in chemistry and biology.  In chemistry spontaneous means thermodynamically favorable, while in standard usage (and generally in biology) spontaneous implies that a reaction is proceeding at a measurable, functionally significant rate.  Yet another insight that emerged through discussions with Melanie Cooper. 

Mike Klymkowsky

Going virtual without a net

Is the coronavirus-based transition from face to face to on-line instruction yet another step to down-grading instructional quality?

It is certainly a strange time in the world of higher education. In response to the current corona virus pandemic, many institutions have quickly, sometimes within hours and primarily by fiat, transitioned from face to face to distance (web-based) instruction. After a little confusion, it appears that laboratory courses are included as well, which certainly makes sense. While virtual laboratories can be built (see our own virtual laboratories in biology)  they typically fail to capture the social setting of a real laboratory.  More to the point, I know of no published studies that have measured the efficacy of such on-line experiences in terms of the ideas and skills students master.

Many instructors (including this one) are being called upon to carry out a radical transformation of instructional practice “on the fly.” Advice is being offered from all sides, from University administrators and technical advisors (see as an example Making Online Teaching a Success).  It is worth noting that much (all?) of this advice falls into the category of “personal empiricism”, suggestions based on various experiences but unsupported  by objective measures of educational outcomes – outcomes that include the extent of student engagement as well as clear descriptions of i) what students are expected to have mastered, ii) what they are expected to be able to do with their knowledge, and iii) what they can actually do. Again, to my knowledge there have been few if any careful comparative studies on learning outcomes achieved via face to face versus virtual teaching experiences. Part of the issue is that many studies on teaching strategies (including recent work on what has been termed “active learning” approaches) have failed to clearly define what exactly is to be learned, a necessary first step in evaluating their efficacy.  Are we talking memorization and recognition, or the ability to identify and apply core and discipline-specific ideas appropriately in novel and complex situations?

At the same time, instructors have not had practical training in using available tools (zoom, in my case) and little in the way of effective support. Even more importantly, there are few published and verified studies to inform what works best in terms of student engagement and learning outcomes. Even if there were clear “rules of thumb” in place to guide the instructor or course designer, there has not been the time or resources needed to implement these changes. The situation is not surprising given that the quality of university level educational programs rarely attracts critical analysis, or the necessary encouragement, support, and recognition needed to make it a departmental priority (see Making education matter in higher education).  It seems to me that the current situation is not unlike attempting to perform a complicated surgery after being told to watch a 3 minute youtube video. Unsurprisingly patient (student learning) outcomes may not be pretty.     

Much of what is missing from on-line instructional scenarios is the human connection, the ability of an instructor to pay attention to how students respond to the ideas presented. Typically this involves reading the facial expressions and body language of students, and through asking challenging (Socratic) questions – questions that address how the information presented can be used to generate plausible explanations or to predict the behavior of a system. These are interactions that are difficult, if not impossible to capture in an on-line setting.

While there is much to be said for active engagement/active learning strategies (see Hake 1998, Freeman et al 2014 and Theobald et al 2020), one can easily argue that all effective learning scenarios involve an instructor who is aware and responsive to students’ pre-existing knowledge. It is also important that the instructor has the willingness (and freedom) to entertain their questions, confusions, and the need for clarification (saying it a different way), or when it may be necessary to revisit important, foundational, ideas and skills – a situation that can necessitate discarding planned materials and “coaching up” students on core concepts and their application. The ability of the instructor to customize instruction “on the fly” is one of the justifications for hiring disciplinary experts in instructional positions, they (presumably) understand the conceptual foundations of the materials they are called upon to present. In its best (Socratic) form, the dialog between student and instructor drives students (and instructors) to develop a more sophisticated and metacognitive understanding of the web of ideas involved in most scientific explanations.

In the absence of an explicit appreciation of the importance of the human interactions between instructor and student, interactions already strained in the context of large enrollment courses, we are likely to find an increase in the forces driving instruction to become more and more about rote knowledge, rather than the higher order skills associated with the ability to juggle ideas, identifying those needed and those irrelevant to a specific situation.  While I have been trying to be less cynical (not a particularly easy task in the modern world), I suspect that the flurry of advice on how to carry out distance learning is more about avoiding the need to refund student fees than about improving students’ educational outcomes (see Colleges Sent Students Home. Now Will They Refund Tuition?)

A short post-script (17 April 2020): Over the last few weeks I have put together the tools to make the on-line MCDB 4650 Developmental Biology course somewhat smoother for me (and hopefully the students). I use Keynote (rather than Powerpoint) for slides; since the iPad is connected wirelessly to the project, this enables me to wander around the class room. The iOS version of Keynote enables me, and students, to draw on slides. Now that I am tethered, I rely more on pre-class beSocratic activities and the Mirroring360 application to connect my iPad to my laptop for Zoom sessions. I am back to being more interactive with the materials presented. I am also starting to pick students at random to answer questions & provide explanations (since they are quiet otherwise) – hopefully that works. Below (↓) is my set up, including a good microphone, laptop, iPad, and the newly arrived volume on Active Learning.

Aggregative and clonal metazoans (a biofundamentalist perspective)

21st Century DEVO-2  In the first post in this series [link], I introduced the observation that single celled organisms can change their behaviors, often in response to social signals.  They can respond to changing environments and can differentiate from one cellular state to the another. Differentiation involves changes in which sets of genes are expressed, which polypeptides and proteins are made [previous post], where the proteins end up within the cell, and which behaviors are displayed by the organism. Differentiation enables individuals to adapt to hostile conditions and to exploit various opportunities. 

The ability of individuals to cooperate with one another, through processes such as quorum sensing, enables them to tune their responses so that they are appropriate and useful. Social interactions also makes it possible for them to produce behaviors that would be difficult or impossible for isolated individuals.  Once individual organisms learn, evolutionarily, how to cooperate, new opportunities and challenges (cheaters) emerge. There are strategies that can enable an organism to adapt to a wider range of environments, or to become highly specialized to a specific environment,  through the production of increasingly complex behaviors.  As described previously, many of these cooperative strategies can be adopted by single celled organisms, but others require a level of multicellularity.  Multicellularity can be transient – a pragmatic response to specific conditions, or it can be (if we ignore the short time that gametes exist as single cells) permanent, allowing the organism to develop the range of specialized cells types needed to build large, macroscopic organisms with complex and coordinated behaviors. In appears that various forms of multicellularity have arisen independently in a range of lineages (Bonner, 1998; Knoll, 2011). We can divide multicellularity into two distinct types, aggregative and clonal – which we will discuss in turn (1).  Aggregative (transient) multicellularity:  Once organisms had developed quorum sensing, they can monitor the density of related organisms in their environment and turn or (or off) specific genes (or sets of genes, necessary to produce a specific behavior.  While there are many variants, one model for such  a behavior is  a genetic toggle switch, in which a particular gene (or genes) can be switched on or off in response to environmental signals acting as allosteric regulators of transcription factor proteins (see Gardner et al., 2000).  Here is an example of an activity (↓) that we will consider in class to assess our understanding of the molecular processes involved.

One outcome of such a signaling system is to provoke the directional migration of amoeba and their aggregation to form the transient multicellular “slug”.  Such behaviors has been observed  in a range of normally unicellular organisms (see Hillmann et al., 2018)(↓). The classic example is  the cellular slime mold Dictyostelium discoideum (Loomis, 2014).  Under normal conditions, these unicellular amoeboid eukaryotes migrate, eating bacteria and such. In this state, the range of an individual’s movement is restricted to short distances.  However when conditions turn hostile, specifically a lack of necessary nitrogen compounds, there is a compelling reason to abandon one environment and migrate to another, more distant that a single-celled organism could reach. This is a behavior that depends upon the presence of a sufficient density (cells/unit volume) of cells that enables them to: 1) recognize one another’s presence (through quorum sensing), 2) find each other through directed (chemotactic) migration, and 3) form a multicellular slug that can go on to differentiate. Upon differentiation about 20% of the cells differentiate (and die), forming a stalk that lifts the other ~80% of the cells into the air.  These non-stalk cells (the survivors) differentiate into spore (resistant to drying out) cells that are released into the air where they can be carried to new locations, establishing new populations.  

The process of cellular differentiation in D. discoideum has been worked out in molecular detail and involves two distinct signaling systems: the secreted pre-starvation factor (PSF) protein and cyclic AMP (cAMP).  PSF is a quorum signaling protein that also serves to activate the cell aggregation/differentiation program (FIG. ↑). If bacteria, that is food, are present, the activity of PSF is inhibited and  cells remain in their single cell state. The key regulator of downstream aggregation and differentiation is the cAMP-dependent protein kinase PKA.  In the unicellular state, PKA activity is inhibited by PufA.  As PSF increases, while food levels decrease, YakA activity increases, inactivating PufA, leading to increased PKA activity.  Active PKA induces the synthesis of two downstream proteins, adenylate cyclase (ACA) and the cAMP receptor (CAR1). ACA catalyzes cAMP synthesis, much of which is secreted from the cell as a signaling molecule. The membrane-bound CAR1 protein acts as a receptor for autocrine (on the cAMP secreting cell) and paracrine (on neighboring cells) signaling.  The binding of cAMP to CAR1 leads to further activation of PKA, increasing cAMP synthesis and secretion – a positive feed-back loop. As cAMP levels increase, downstream genes are activated (and inhibited) leading cells to migrate toward one another, their adhesion to form a slug.  Once the slug forms and migrates to an appropriate site, the process of differentiation (and death) leading to stalk and spore formation begins. The fates of the aggregated cells is determined stochastically, but social cheaters can arise. Mutations can lead to individuals that avoid becoming stalk cells.  In the long run, if all individuals were to become cheaters, it would be impossible to form a stalk, so the purpose of social cooperation would be impossible to achieve.  In the face of environmental variation, populations invaded by cheaters are more likely to become extinct.  For our purposes the various defenses against cheaters are best left to other courses (see here if interested Strassmann et al., 2000).  

Clonal (permanent) multicellularity:  The type of multicellularity that most developmental biology courses focus on is what is termed clonal multicellularity – the organism is a clone of an original cell, the zygote, a diploid cell produced by the fusion of sperm and egg, haploid cells formed through the process of meiosis (2).  It is during meiosis that most basic genetic processes occur, that is the recombination between maternal and paternal chromosomes leading to the shuffling of alleles along a chromosome, and the independent segregation of chromosomes to form haploid gametes, gametes that are genetically distinct from those present in either parent. Once the zygote forms, subsequent cell divisions involve mitosis, with only a subset of differentiated cells, the cells of the germ line, capable of entering meiosis.  

Non-germ line, that is somatic cells, grow and divide. They interact with one another directly and through various signaling processes to produce cells with distinct patterns of gene expression, and so differentiated behaviors.  A key difference from a unicellular organism, is that the cells will (largely) stay attached to one another, or to extracellular matrix materials secreted by themselves and their neighbors.  The result is ensembles of cells displaying different specializations and behaviors.  As such cellular colonies get larger, they face a number of physical constraints – for example, cells are open non-equilibrium systems, to maintain themselves and to grow and reproduce, they need to import matter and energy from the external world. Cells also produce a range of, often toxic, waste products that need to be removed.  As the cluster of zygote-derived cells grows larger, and includes more and more cells, some cells will become internal and so cut off from necessary resources. While diffusive processes are often adequate when a cell is bathed in an aqueous solution, they are inadequate for a cell in the interior of a large cell aggregate (3).  The limits of diffusive processes necessitate other strategies for resource delivery and waste removal; this includes the formation of tubular vascular systems (such as capillaries, arteries, veins) and contractile systems (hearts and such) to pump fluids through these vessels, as well as cells specialized to process and transport a range of nutrients (such as blood cells).  As organisms get larger, their movements require contractile machines (muscle, cartilage, tendons, bones, etc) driving tails, fins, legs, wings, etc. The coordination of such motile systems involves neurons, ganglia, and brains. There is also a need to establish barriers between the insides of an organism and the outside world (skin, pulmonary, and gastrointestinal linings) and the need to protect the interior environment from invading pathogens (the immune system).  The process of developing these various systems depends upon controlling patterns of cell growth, division, and specialization (consider the formation of an arm), as well as the controlled elimination of cells (apoptosis), important in morphogenesis (forming fingers from paddle-shaped appendages), the maturation of the immune system (eliminating cells that react against self), and the wiring up, and adaptation of the nervous system. Such changes are analogous to those involved in aggregative multicellularity.      

Origins of multicellularity:  While aggregative multicellularity involves an extension of quorum sensing and social cooperation between genetically distinct, but related individuals, we can wonder whether similar drivers are responsible for clonal multicellularity.  There are a number of imaginable adaptive (evolutionary) drivers but two spring to mind: a way to avoid predators by getting bigger than the predators and as a way to produce varied structures needed to exploit various ecological niches and life styles. An example of the first type of driver of multicellularity is offered by the studies of Boraas et al  (1998). They cultured the unicellular green alga Chlorella vulgaris, together with a unicellular predator, the phagotrophic flagellated protist Ochromonas vallescia. After less than 100 generations (cell divisions), they observed the appearance of multicellular, and presumable inedible (or at least less easily edible), forms. Once selected, this trait appears to be stable, such that “colonies retained the eight-celled form indefinitely in continuous culture”.  To my knowledge, the genetic basis for this multicellularity remains to be determined.  

Cell Differentiation:  One feature of simple colonial organisms is that when dissociated into individual cells, each cell is capable of regenerating a new organism. The presence of multiple (closely related) cells in a single colony opens up the possibility of social interactions; this is distinct from the case in aggregative multicellularity, where social cooperation came first. Social cooperation within a clonal metazoan means that most cells “give up” their ability to reproduce a new organism (a process involving meiosis). Such irreversible social interactions mark the transition from a colonial organism to a true multicellular organism. As social integration increases, cells can differentiate so as to perform increasingly specialized functions, functions incompatible with cell division. Think for a moment about a human neuron or skeletal muscle cell – in both cases, cell division is no longer possible (apparently). Nevertheless, the normal functioning of such cells enhances the reproductive success of the organism as a whole – a classic example of inclusive fitness (remember heterocysts?)  Modern techniques of single cell sequencing and data analysis have now been employed to map this process of cellular differentiation in increasingly great detail, observations that will inform our later discussions (see Briggs et al., 2018 and future posts). In contrast, the unregulated growth of a cancer cell is an example of an asocial behavior, an asocial behavior that is ultimately futile, except in those rare cases (four known at this point) in which a cancer cell can move from one organism to another (Ujvari et al., 2016).  

Unicellular affordances for multicellularity:  When considering the design of a developmental biology course, we are faced with the diversity of living organisms – the basic observation that Darwin, Wallace, their progenitors and disciplinary descendants set out to solve. After all there are many millions of different types of organisms; among the multicellular eukaryotes, there are six major group : the ascomycetes and basidiomycetes fungi, the florideophyte red algae, laminarialean brown algae, embryophytic land plants and animals (Knoll, 2011 ↑).  Our focus will be on animals. “All members of Animalia are multicellular, and all are heterotrophs (i.e., they rely directly or indirectly on other organisms for their nourishment). Most ingest food and digest it in an internal cavity.” [Mayer link].  From a macroscopic perspective, most animals have (or had at one time during their development) an anterior to posterior, that is head to tail, axis. Those that can crawl, swim, walk, or fly typically have a dorsal-ventral or back to belly axis, and some have a left-right axis as well.  

But to be clear, a discussion of the various types of animals is well beyond the scope of any introductory course in developmental biology, in part because there are 35 (assuming no more are discovered) different “types” (phyla) of animals – nicely illustrated at this website [BBC: 35 types of animals, most of whom are really weird)].  So again, our primary focus will be on one group, the vertebrates – humans are members of this group.  We will also consider experimental insights derived from studies of various “model” systems, including organisms from another metazoan group, the  ecdysozoa (organisms that shed their outer layer as they grow bigger), a group that includes fruit flies and nematode worms. 

My goal will be to ignore most of the specialized terminology found in the scholarly literature, which can rapidly turn a biology course into a vocabulary lesson and that add little to understanding of basic processes relevant to a general understanding of developmental processes (and relevant to human biology, medicine, and biotechnology). This approach is made possible by the discovery that the basic processes associated with animal (and metazoan) development are conserved. In this light, no observation has been more impactful than the discovery that the nature and organization of the genes involved in specifying the head to tail axes of the fruit fly and vertebrates (such as the mouse and human) is extremely similar in terms of genomic organization and function (Lappin et al., 2006 ↓), an observation that we will return to repeatedly.  Such molecular similarities extend to cell-cell and cell-matrix adhesion systems, systems that release and respond to various signaling molecules, controlling cell behavior and gene expression, and reflects the evolutionary conservation and the common ancestry of all animals (Brunet and King, 2017; Knoll, 2011). 

What can we know about the common ancestor of the animals?  Early on in the history of comparative cellular anatomy, the striking structural similarities between  the feeding system of choanoflagellate protozoans, a motile (microtubule-based) flagellum a surrounded by a “collar”of microfilament-based microvilli) and a structurally similar organelle in a range of multicellular organisms led to the suggestion that choanoflagellates and animals shared a common ancestor.  The advent of genomic sequencing and analysis has only strengthened this hypothesis, namely that choanoflagellates and animals form a unified evolutionary clade, the ‘Choanozoa’  (see tree↑ above)(Brunet and King, 2017).  Moreover, “many genes required for animal multicellularity (e.g., tyrosine kinases, cadherins, integrins, and extracellular matrix domains) evolved before animal origins”.  The implications is that the Choanozoan ancestor was predisposed to exploit some of the early opportunities offered by clonal multicellularity. These pre-existing affordances, together with newly arising genes and proteins (Long et al., 2013) were exploited in multiple lineages in the generation of multicellular organisms (see Knoll, 2011).

Basically to understand what happened next, some ~600 million years ago or so, we will approach the various processes involved in the shaping of animal development.  Because all types of developmental processes, including the unicellular to colonial transition, involve changes in gene expression, we will begin with the factors involved in the regulation of gene expression.  

1). Please excuse the inclusive plural, but it seems appropriate in the context of what I hope will be a highly interactive course.
2). I will explicitly ignore variants as (largely) distractions, better suited for more highly specialized courses.
3). We will return to this problem when (late in the course, I think) we will discuss the properties of induced pluripotent stem cell (iPSC) derived organoids.

On teaching developmental biology in the 21st century (a biofundamentalist perspective)

On teaching developmental biology and trying to decide where to start: differentiation

Having considered the content of courses in chemistry [1] and  biology [2, 3], and preparing to teach developmental biology for the first time, I find myself reflecting on how such courses might be better organized.  In my department, developmental biology (DEVO) has returned after a hiatus as the final capstone course in our required course sequence, and so offers an opportunity within which to examine what students have mastered as they head into their more specialized (personal) educational choices.  Rather than describe the design of the course that I will be teaching, since at this point I am not completely sure what will emerge, what I intend to do (in a series of posts) is to describe, topic by topic, the progression of key concepts, the observations upon which they are based, and the logic behind their inclusion.

Modern developmental biology emerged during the mid-1800s from comparative embryology [4] and was shaped by the new cell theory (the continuity of life and the fact that all organisms are composed of cells and their products) and the ability of cells to differentiate, that is, to adopt different structures and behaviors [5].  Evolutionary theory was also key.  The role of genetic variation based on mutations and selection, in the generation of divergent species from common ancestors, explained why a single, inter-connected Linnaean (hierarchical) classification system (the phylogenic tree of life →) of organisms was possible and suggested that developmental mechanisms were related to similar processes found in their various ancestors. 

So then, what exactly are the primary concepts behind developmental biology and how do they emerge from evolutionary, cell, and molecular biology?  The concept of “development” applies to any process characterized by directional changes over time.  The simplest such process would involve the progress from the end of one cell division event to the beginning of the next; cell division events provide a convenient benchmark.  In asexual species, the process is clonal, a single parent gives rise to a genetically identical (except for the occurrence of new mutations) offspring. Often there is little distinction between parent and offspring.  In sexual species, a dramatic and unambiguous benchmark involves the generation of a new and genetically distinct organism.  This “birth” event is marked by the fusion of two gametes (fertilization) to form a new diploid organism.  Typically gametes are produced by a complex cellular differentiation process (gametogenesis), ending with meiosis and the formation of haploid cells.  In multicellular organisms, it is often the case that a specific lineage of cells (which reproduce asexually), known as the germ line, produce the gametes.  The rest of the organism, the cells that do not produce gametes, is known as the soma, composed of somatic cells.   Cellular continuity remains, however, since gametes are living (albeit haploid) cells.  

It is common for the gametes that fuse to be of two different types, termed oocyte and sperm.  The larger, and generally immotile gamete type is called an oocyte and an individual that produces oocytes is termed female. The smaller, and generally motile gamete type is called a sperm; individuals that produces sperm are termed male. Where a single organism can produce both oocytes and sperm, either at the same time or sequentially, they are referred to as hermaphrodites (named after Greek Gods, the male Hermes and the female Aphrodite). Oocytes and sperm are specialized cells; their formation involves the differential expression of genes and the specific molecular mechanisms that generate the features characteristic of the two cell types.  The fusion of gametes, fertilization,  leads to a zygote, a diploid cell that (usually) develops into a new, sexually mature organism.    

An important feature of the process of fertilization is that it requires a level of social interaction, the two fusing cells (gametes) must recognize and fuse with one another.  The organisms that produce these gametes must cooperate; they need to produce gametes at the appropriate time and deliver them in such a way that they can find and recognize each other and avoid “inappropriate” interactions”.  The specificity of such interactions underlie the reproductive isolation that distinguishes one species from another.  The development of reproductive isolation emerges as an ancestral population of organisms diverges to form one or more new species.  As we will see, social interactions, and subsequent evolutionary effects, are common in the biological world.  

The cellular and molecular aspects of development involve the processes by which cells grow, replicate their genetic material (DNA replication), divide to form distinct parent-offspring or similar sibling cells, and may alter their morphology (shape), internal organization, motility, and other behaviors, such as the synthesis and secretion of various molecules, and how these cells respond to molecules released by other cells.  Developmental processes involve the expression and the control of all of these processes.

Essentially all changes in cellular behavior are associated with changes in the activities of biological molecules and the expression of genes, initiated in response to various external signaling events – fertilization itself is such a signal.  These signals set off a cascade of regulatory interactions, often leading to multiple “cell types”, specialized for specific functions (such as muscle contraction, neural and/or hormonal signaling, nutrient transport, processing, and synthesis,  etc.).  For specific parts of the organism, external or internal signals can result in a short term “adaptive” response (such as sweating or panting in response to increased internal body temperature), after which the system returns to its original state, or in the case of developing systems, to new states, characterized by stable changes in gene expression, cellular morphology, and behavior.    

Development in bacteria (and other unicellular organisms):  In most unicellular organisms, the cell division process is reasonably uneventful, the cells produced are similar to the original cell – but not always.  A well studied example is the bacterium Caulobacter crescentus (and related species) [link][link].  In cases such as this, the process of growth  leads to phenotypically different daughters.  While it makes no sense to talk about a beginning (given the continuity of life after the appearance of the last universal common ancestor or LUCA), we can start with a “swarmer” cell, characterized by the presence of a motile flagellum (a molecular machine driven by coupled chemical reactions – see past blogpost] that drives motility [figure modified from 6 ]. 

A swarmer will eventually settle down, loose the flagellum, and replace it with a specialized structure (a holdfast) designed to anchor the cell to a solid substrate.  As the organism grows, the holdfast develops a stalk that lifts the cell away from the substrate.  As growth continues, the end of the cell opposite the holdfast begins to differentiate (becomes different) from the holdfast end of the cell – it begins the process leading to the assembly of a new flagellar apparatus.  When reproduction (cell growth, DNA replication, and cell division) occurs, a swarmer cell is released and can swim away and colonize another area, or settle nearby.  The holdfast-anchored cell continues to grow, producing new swarmers.  This process is based on the inherent asymmetry of the system – the holdfast end of the cell is molecularly distinct from the flagellar end [see 7].

The process of swarmer cell formation in Caulobacter is an example of what we will term deterministic phenotypic switching.  Cells can also exploit molecular level noise (stochastic processes) that influence gene expression to generate phenotypic heterogeneity, different behaviors expressed by genetically identical cells within the same environment [see 8, 9].  Molecular noise arises from the random nature of molecular movements and the rather small (compared to macroscopic systems) numbers of most molecules within a cell.  Most cells contain one or two copies of any particular gene, and a similarly small number of molecular sequences involved in their regulation [10].  Which molecules are bound to which regulatory sequence, and for how long, is governed by inter-molecular surface interactions and thermally driven collisions, and is inherently noisy.  There are strategies that can suppress but not eliminate such noise [see 11].  As dramatically illustrated by Elowitz  and colleagues [8](), molecular level noise can produce cells with different phenotypes.  Similar processes are active in eukaryotes (including humans), and can lead to the expression of one of the two copies of a gene (mono-allelic expression) present in a diploid organism.  This can lead to effects such as haploinsufficiency and selective (evolutionary) lineage effects if the two alleles are not identical [12, 13]. Such phenotypic heterogeneity among what are often genetically identical cells is a topic that is rarely discussed (as far as I can discern) in introductory cell, molecular, or developmental biology courses [past blogpost].

The ability to switch phenotypes can be a valuable trait if an organism’s environment is subject to significant changes.  As an example, when the environment gets hostile, some bacterial cells transition from a rapidly dividing to a slow or non-dividing state.  Such “spores” can differentiate so as to render them highly resistant to dehydration and other stresses.  If changes in environment are very rapid, a population can protect itself by continually having some cells (stochastically) differentiating into spores, while others continue to divide rapidly. Only a few individuals (spores) need to survive a catastrophic environmental change to quickly re-establish the population.

Dying for others – social interactions between “unicellular” organisms:  Many students might not predict that one bacterial cell would “sacrifice” itself for the well being of others, but in fact there are a number of examples of this type of self-sacrificing behavior, known as programmed cell death, which is often a stochastic process.  An interesting example is provided by cellular specialization for photosynthesis or nitrogen fixation in cyanobacteria [see 9].  These two functions require mutually exclusive cellular environments to occur, in particular the molecular oxygen (O2) released by photosynthesis inhibits the process of nitrogen fixation.  Nevertheless, both are required for optimal growth.  The solution?  some cells differentiate into what are known as heterocysts, cells committed to nitrogen fixation ( a heterocyst in Anabaena spiroides, adapted from link), while most ”vegetative” cells continue with photosynthesis.  Heterocysts cannot divide, and eventually die – they sacrifice themselves for the benefit of their neighbors, the vegetative cells, cells that can reproduce.

The process by which the death of an individual can contribute resources that can be used to insure or enhance the survival and reproduction of surrounding individuals is an inherently social process, and is subject of social evolutionary mechanisms [14, 15][past blogpost].  Social behaviors can be selected for because the organism’s neighbors, the beneficiaries of their self-sacrifice are likely to be closely (clonally) related to themselves.  One result of the social behavior is, at the population level, an increase in one aspect of evolutionary fitness,  termed “inclusive fitness.”  

Such social behaviors can enable a subset of the population to survive various forms of environmental stress (see spore formation above).  An obvious environmental stress involves the impact of viral infection.  Recall that viruses are completely dependent upon the metabolic machinery of the infected cell to replicate. While there are a number of viral strategies, a common one is bacterial lysis – the virus replicates explosively, kills the infected cells, leading to the release of virus into the environment to infect others.  But, what if the infected cell kills itself BEFORE the virus replicates – the dying (self-sacrificing, altruistic) cell “kills” the virus (although viruses are not really alive) and stops the spread of the infection.  Typically such genetically programmed cell death responses are based on a simple two-part system, involving a long lived toxin and a short-lived anti-toxin.  When the cell is stressed, for example early during viral infection, the level of the anti-toxin can fall, leading to the activation of  the toxin. 

Other types of social behavior and community coordination (quorum effects):  Some types of behaviors only make sense when the density of organisms rises above a certain critical level.  For example,  it would make no sense for an Anabaena cell  to differentiate into a heterocyst (see above) if there are no vegetative cells nearby.  Similarly, there are processes in which a behavior of a single bacterial cell, such as the synthesis and secretion of a specific enzyme, a specific import or export machine,  or the construction of a complex, such as a DNA uptake machine, makes no sense in isolation – the secreted molecule will just diffuse away, and so be ineffective, the molecule to be imported (e.g. lactose) or exported (an antibiotic) may not be present, or there may be no free DNA to import.  However, as the concentration (organisms per volume) of bacteria increases, these behaviors can begin to make biological sense – there is DNA to eat or incorporate and the concentration of secreted enzyme can be high enough to degrade the target molecules (so they are inactivated or can be imported as food).   

So how does a bacterium determine whether it has neighbors or whether it wants to join a community of similar organisms?  After all, it does not have eyes to see. The process used is known as quorum sensing.  Each individual synthesizes and secretes a signaling molecule and a receptor protein whose activity is regulated by the binding of the signaling molecule.  Species specificity in signaling molecules and receptors insures that organisms of the same kind are talking to one another and not to other, distinct types of organisms that may be in the environment.   At low signaling molecule concentrations, such as those produced by a single bacterium in isolation, the receptor is not activated, and the cell’s behavior remains unchanged.  However, as the concentration of bacteria increases, the concentration of the signal increases, leading to the activation of the receptor.  Activation of the receptor can have a number of effects, including increased synthesis of the signal and other changes, such as movement in response to signal s- through regulation of flagellar and other motility systems, such a system can lead to the directed migration (aggregation) of cells [see 16].   

In addition to driving the synthesis of a common good (such as a useful extracellular molecule), social interactions can control processes such as  programmed cell death.  When the concentration of related neighbors is high, the programmed death of an individual can be beneficial, it can  lead to release of nutrients,  common goods, including DNA molecules, that can be used by neighbors (relatives)[17, 18] – an increase in the probability of cell death in response to a quorum can increased in a way that increases inclusive fitness.  On the other hand,  if there are few related individuals in the neighborhood, programmed cell death “wastes” these resources, and so is likely to be suppressed (you might be able to generate plausible mechanisms that could control the probability of programmed cell death).     

As we mentioned previously with respect to spore formation, the generation of a certain percentage of “persisters” – individuals that withdraw from active growth and cell division, can enable a population to survive stressful situations, such as the presence of an antibiotic.  On the other hand, generating too many persisters may place the population at a reproductive disadvantage.  Once the antibiotic is gone, the persisters can return into active division. The ability of bacteria to generate persisters is a serious problem in treating people with infections [19].  

Of course, as in any social system, the presumption of cooperation (expending energy to synthesize the signal, sacrificing oneself for others) can open the system to cheaters [blogpost].  All such “altruistic” behaviors are vulnerable to cheaters.(1)  For example, a cheater that avoids programmed cell death (for example due to an inactivating mutation that effects the toxin molecule involved) will come to take over the population.  The downside, for the population, is that if cheaters take over,  the population is less likely to survive the environmental events that the social behavior was evolve to address.  In response to the realities of cheating, social organisms need to adopt various social-validation systems [see 20 as an example]; we see this pattern of social cooperation, cheating, and social defense mechanism throughout the biological world. 

Follow-on posts:


1): Such as people who fail to pay their taxes or disclose their tax returns.

Genes – way weirder than you thought

Pretty much everyone, at least in societies with access to public education or exposure to media in its various forms, has been introduced to the idea of the gene, but “exposure does not equate to understanding” (see Lanie et al., 2004).  Here I will argue that part of the problem is that instruction in genetics (or in more modern terms, the molecular biology of the gene and its role in biological processes) has not kept up with the advances in our understanding of the molecular mechanisms underlying biological processes (Gayon, 2016).  

Let us reflect (for a moment) on the development of the concept of a gene: Over the course of human history, those who have been paying attention to such things have noticed that organisms appear to come in “types”, what biologists refer to as species. At the same time, individual organisms of the same type are not identical to one  another, they vary in various ways. Moreover, these differences can be passed from generation to generation, and by controlling  which organisms were bred together; some of the resulting offspring often displayed more extreme versions of the “selected” traits.  By strictly controlling which individuals were bred together, over a number of generations, people were able to select for the specific traits they desired (→).  As an interesting aside, as people domesticated animals, such as cows and goats, the availability of associated resources (e.g. milk) led to reciprocal effects – resulting in traits such as adult lactose tolerance (see Evolution of (adult) lactose tolerance & Gerbault et al., 2011).  Overall, the process of plant and animal breeding is generally rather harsh (something that the fanciers of strange breeds who object to GMOs might reflect upon), in that individuals that did not display the desired trait(s) were generally destroyed (or at best, not allowed to breed).  

Charles Darwin took inspiration from this process, substituting “natural” for artificial (human-determined) selection to shape populations, eventually generating new species (Darwin, 1859).  Underlying such evolutionary processes was the presumption that traits, and their variation, was “encoded” in some type of “factors”, eventually known as genes and their variants, alleles.  Genes influenced the organism’s molecular, cellular, and developmental systems, but the nature of these inheritable factors and the molecular trait building machines active in living systems was more or less completely obscure. 

Through his studies on peas, Gregor Mendel was the first to clearly identify some of the rules for the behavior of these inheritable factors using highly stereotyped, and essentially discontinuous traits – a pea was either yellow or green, wrinkled or smooth.  Such traits, while they exist in other organisms, are in fact rare – an example of how the scientific exploration of exceptional situations can help understand general processes, but the downside is the promulgation of the idea that genes and traits are somehow discontinuous – that a trait is yes/no, displayed by an organism or not – in contrast to the realities that the link between the two is complex, a reality rarely directly addressed (apparently) in most introductory genetics courses.  Understanding such processes is critical to appreciating the fact that genetics is often not destiny, but rather alterations in probabilities (see Cooper et al., 2013).  Without such an more nuanced and realistic understanding, it can be difficult to make sense of genetic information.      

A gene is part of a molecular machine:  A number of observations transformed the abstraction of Darwin’s and Mendel’s hereditary factors into physical entities and molecular mechanisms (1).  In 1928 Fred Griffith demonstrated that a genetic trait could be transferred from dead to living organisms – implying a degree of physical / chemical stability; subsequent observations implied that the genetic information transferred involved DNA molecules. The determination of the structure of double-stranded DNA immediately suggested how information could be stored in DNA (in variations of bases along the length of the molecule) and how this information could be duplicated (based on the specificity of base pairing).  Mutations could be understood as changes in the sequence of bases along a DNA molecule (introduced by chemicals, radiation, mistakes during replication, or molecular reorganizations associated with DNA repair mechanisms and selfish genetic elements.  

But on their own, DNA molecules are inert – they have functions only within the context of a living organism (or highly artificial, that is man made, experimental systems).  The next critical step was to understand how a gene works within a biological system, that is, within an organism.  This involve appreciating the molecular mechanisms (primarily proteins) involved in identifying which stretches of a particular DNA molecule were used as templates for the synthesis of RNA molecules, which in turn could be used to direct the synthesis of polypeptides (see previous post on polypeptides and proteins).  In the context of the introductory biology courses I am familiar with (please let me know if I am wrong), these processes are based on a rather deterministic context; a gene is either on or off in a particular cell type, leading to the presence or absence of a trait. Such a deterministic presentation ignores the stochastic nature of molecular level processes (see past post: Biology education in the light of single cell/molecule studies) and the dynamic interaction networks that underlie cellular behaviors.   

But our level of resolution is changing rapidly (2).  For a number of practical reasons, when the human genome was first sequence, the identification of polypeptide-encoding genes was based on recognizing “open-reading frames” (ORFs) encoding polypeptides of > 100 amino acids in length (> 300 base long coding sequence).  The increasing sensitivity of mass spectrometry-based proteomic studies reveals that smaller ORFs (smORFs) are present and can lead to the synthesis of short (< 50 amino acid long) polypeptides (Chugunova et al., 2017; Couso, 2015).  Typically an ORF was considered a single entity – basically one gene one ORF one polypeptide (3).  A recent, rather surprising discovery is what are known as “alternative ORFs” or altORFs; these RNA molecules that use alternative reading frames to encode small polypeptides.  Such altORFs can be located upstream, downstream, or within the previously identified conventional ORF (figure →)(see Samandi et al., 2017).  The implication, particularly for the analysis of how variations in genes link to traits, is that a change, a mutation or even the  experimental  deletion of a gene, a common approach in a range of experimental studies, can do much more than previously presumed – not only is the targeted ORF effected, but various altORFs can also be modified.  

The situation is further complicated when the established rules of using RNAs to direct polypeptide synthesis via the process of translation, are violated, as occurs in what is known as “repeat-associated non-ATG (RAN)” polypeptide synthesis (see Cleary and Ranum, 2017).  In this situation, the normal signal for the start of RNA-directed polypeptide synthesis, an AUG codon, is subverted – other RNA synthesis start sites are used leading to underlying or imbedded gene expression.  This process has been found associated with a class of human genetic diseases, such as amyotrophic lateral sclerosis (ALS) and frontotemporal dementia (FTD) characterized by the expansion of simple (repeated) DNA sequences  (see Pattamatta et al., 2018).  Once they exceed a certain length, such“repeat” regions have been found to be associated with the (apparently) inappropriate transcription of RNA in both directions, that is using both DNA strands as templates (← A: normal situation, B: upon expansion of the repeat domain).  These abnormal repeat region RNAs are translated via the RAN process to generate six different types of toxic polypeptides. 

So what are the molecular factors that control the various types of altORF transcription and translation?  In the case of ALS and FTD, it appears that other genes, and the polypeptides and proteins they encode, are involved in regulating the expression of repeat associated RNAs (Kramer et al., 2016)(Cheng et al., 2018).  Similar or distinct mechanisms may be involved in other  neurodegenerative diseases  (Cavallieri et al., 2017).  

So how should all of these molecular details (and it is likely that there are more to be discovered) influence how genes are presented to students?  I would argue that DNA should be presented as a substrate upon which various molecular mechanisms occur; these include transcription in its various forms (directed and noisy), as well as DNA synthesis, modification, and repair mechanisms occur.   Genes are not static objects, but key parts of dynamic systems.  This may be one reason that classical genetics, that is genes presented within a simple Mendelian (gene to trait) framework, should be moved deeper into the curriculum, where students have the background in molecular mechanisms needed to appreciate its complexities, complexities that arise from the multiple molecular machines acting to access, modify, and use the information captured in DNA (through evolutionary processes), thereby placing the gene in a more realistic cellular perspective (4). 


1. Described greater detail in biofundamentals™

2. For this discussion, I am completely ignoring the roles of genes that encode RNAs that, as far as is currently know, do not encode polypeptides.  That said, as we go on, you will see that it is possible that some such non-coding RNA may encode small polypeptides.  

3. I am ignoring the complexities associated with alternative promoter elements, introns, and the alternative and often cell-type specific regulated splicing of RNAs, to create multiple ORFs from a single gene.  

4. With respects to Norm Pace – assuming that I have the handedness of the DNA molecules wrong or have exchanged Z for A or B. 

When is a gene product a protein when is it a polypeptide?

As a new assistant professor (1), I was called upon to teach my department’s “Cell Biology” course. I found,and still find, the prospect challenging in part because I am not exactly sure which aspects of cell biology are important for students to know, both in the context of the major, as well as their lives and subsequent careers.  While it seems possible (at least to me) to lay out a coherent conceptual foundation for biology as a whole [see 1], cell biology can often appear to students as an un-unified hodge-podge of terms and disconnected cellular systems, topics too often experienced as a vocabulary lesson, rather than as a compelling narrative. As such, I am afraid that the typical cell biology course often re-enforces an all too common view of biology as a discipline, a view, while wrong in most possible ways, was summarized by the 19th/early 20th century physicist Ernest Rutherford as “All science is either physics or stamp collecting.”  A key motivator for the biofundamentals project [2] has been to explore how to best dispel this prejudice, and how to more effectively present to students a coherent narrative and the key foundational observations and ideas by which to scientifically consider living systems, by any measure the most complex systems in the Universe, systems shaped, but not determined by, physical chemical properties and constraints, together with the historical vagaries of evolutionary processes on an ever-changing Earth. 

Two types of information:  There is an underlying dichotomy within biological systems: there is the hereditary information encoded in the sequence of nucleotides along double-stranded DNA molecules (genes and chromosomes).  There is also the information inherent in the living system.  The information in DNA is meaningful only in the context of the living cell, a reaction system that has been running without interruption since the origin of life.  While these two systems are inextricably interconnected, there is a basic difference between them. Cellular systems are fragile, once dead there is no coming back.  In contrast the information in DNA can survive death – it can move from cell to cell in the process of horizontal gene transfer.  The Venter group has replaced the DNA of bacterial cells with synthetic genomes in an effort to define the minimal number of genes needed to support life, at least in a laboratory setting [see 3, 4].  In eukaryotes, cloning is carried out by replacing a cell’s DNA, with that of another cell (reference).  

Conflating protein synthesis and folding with assembly and function: Much of the information stored in a cell’s DNA is used to encode the sequence of various amino acid polymers (polypeptides).  While over-simplified [see 5], students are generally presented with the view that each gene encodes a particular protein through DNA-directed RNA synthesis (transcription) and RNA-directed polypeptide synthesis (translation).  As the newly synthesized polypeptide emerges from the ribosomal tunnel, it begins to fold, and is released into the cytoplasm or inserted into or through a cellular membrane, where it often interacts with one or more other polypeptides to form a protein  [see 6].  The assembled protein is either functional or becomes functional after association with various non-polypeptide co-factors or post-translational modifications.  It is the functional aspect of proteins that is critical, but too often their assembly dynamics are overlooked in the presentation of gene expression/protein synthesis, which is really a combination of distinct processes. 

Students are generally introduced to protein synthesis through the terms primary, secondary, tertiary, and quaternary structure, an approach that can be confusing since many (most) polypeptides are not proteins and many proteins are parts of complex molecular machines [here is the original biofundamentals web page on proteins + a short video][see Teaching without a Textbook]. Consider the nuclear pore complex, a molecular machine that mediates the movement of molecules into and out of the nucleus.  A nuclear pore is “composed of ∼500, mainly evolutionarily conserved, individual protein molecules that are collectively known as nucleoporins (Nups)” [7]. But what is the function of a particular NUP, particularly if it does not exist in significant numbers outside of a nuclear pore?  Is a nuclear pore one protein?  In contrast, the membrane bound, mitochondrial ATP synthase found in aerobic bacteria and eukaryotic mitochondria, is described as composed “of two functional domains, F1 and Fo. F1 comprises 5 different subunits (three α, three β, and one γ, δ and ε)” while “Fo contains subunits c, a, b, d, F6, OSCP and the accessory subunits e, f, g and A6L” [8].  Are these proteins or subunits? is the ATP synthase a protein or a protein complex?  

Such confusions arise, at least in part, from the primary-quaternary view of protein structure, since the same terms are applied, generally without clarifying distinction, to both polypeptides and proteins. These terms emerged historically. The purification of a protein was based on its activity, which can only be measured for an intact protein. The primary structure of  a polypeptide was based on the recognition that DNA-encoded amino acid polymers are unbranched, with a defined sequence of amino acid residues (see Sanger. The chemistry of insulin).  The idea of a polypeptide’s secondary structure was based on the “important constraint that all six atoms of the amide (or peptide) group, which joins each amino acid residue to the next in the protein chain, lie in a single plane” [9], which led Pauling, Corey and Branson [10] to recognized the α-helix and β-sheet, as common structural motifs.  When a protein is composed of a single polypeptide, the final folding pattern of the polypeptide, is referred to as its tertiary structure and is apparent in the first protein structure solved, that of myoglobin (↓), by Max Perutz and John Kendell.  Myoglobin’s role in O2 transport depends upon a non-polypeptide (prosthetic) heme group. So far so good, a gene encodes a polypeptide and as it folds a polypeptide becomes a protein – nice and simple (2).  Complications arise from the observations that 1) many proteins are composed of multiple polypeptides, encoded for by one or more genes, and 2) some polypeptides are a part of different proteins.  Hemoglobin, the second protein whose structure was determined, illustrates the point (←).  Hemoglobin is composed of four polypeptides encoded by distinct genes encoding α- and β-globin polypeptides.  These polypeptides are related in structure, function, and evolutionary origins to myoglobin, as well  as the cytoglobin and neuroglobin proteins (→).  In humans, there are a number of distinct α-like globin and β-like globin genes that are expressed in different hematopoetic tissues during development, so functional hemoglobin proteins can have a number of distinct (albeit similar) subunit compositions and distinct properties, such as their affinities for O2 [see 11].  

But the situation often gets more complicated.  Consider centrin-2, a eukaryotic Ca2+ binding polypeptide that plays roles in organizing microtubules, building cilia, DNA repair, and gene expression [see 12 and references therein].  So, is the centrin-2 polypeptide just a polypeptide, a protein, or a part of a number of other proteins?  As another example, consider the basic-helix-loop-helix family of transcription factors; these transcription factor proteins are typically homo- or hetero-dimeric; are these polypeptides proteins in their own right?  The activity of these transcription factors is regulated in part by which binding partners they contain. bHLH polypeptides also interact with the Id polypeptide (or is it a protein); Id lacks a DNA binding domain so when it forms a dimer with a bHLH polypeptide it inhibits DNA binding (↓).  So is a single bHLH polypeptide a protein or is the protein necessarily a dimer?  More to the point, does the current primary→quaternary view of protein structure help or hinder student understanding of the realities of biological systems?  A potentially interesting bio-education research question.

A recommendation or two:  While under no illusion that the complexities of polypeptide synthesis and protein assembly can be easily resolved – it is surely possible to present them in a more coherent, consistent, and accessible manner.  Here are a few suggestions that might provoke discussion.  Let us first recognize that, for those genes that encode polypeptides: i) they encode polypeptides rather than functional proteins (a reality confused by the term “quaternary structure”).  We might well distinguish a polypeptide from a protein based on the concentration of free monomeric polypeptide (gene product) within the cell.  Then we need to convey the reality to students that the assembly of a protein is no simple process, particularly within the crowded cytoplasm [13], a misconception supported by the simple secondary-tertiary structure perspective. While some proteins assemble on their own, many (most?) cannot.

As an example, consider the protein tubulin (↑). As noted by Nithianantham et al [14], “ Five conserved tubulin cofactors/chaperones and the Arl2 GTPase regulate α- and β-tubulin assembly into heterodimers” and the “tubulin cofactors TBCD, TBCE, and Arl2, which together assemble a GTP-hydrolyzing tubulin chaperone critical for the biogenesis, maintenance, and degradation of soluble αβ-tubulin.”  Without these various chaperones the tubulin protein cannot be formed.  Here the distinction between protein and multiprotein complex is clear, since tubulin protein exists in readily detectable levels within the cell, in contrast to the α- and β-tubulin polypeptides, which are found complexed to the TBCB and TBCA chaperone polypeptides. Of course the balance between tubulin and tubulin polymers (microtubules) is itself regulated by a number of factors. 

 The situation is even more complex when we come to the ribosome and other structures, such as the nuclear pore.  Woolford [15] estimates that “more than 350 protein and RNA molecules participate in yeast ribosome assembly, and many more in metazoa”; in addition to four ribsomal RNAs and ~80 polypeptides (often referred to as ribosomal proteins) that are synthesized in the cytoplasm and transported into the nucleus in association with various transport factors, these “assembly factors, including diverse RNA-binding proteins, endo- and exonucleases, RNA helicases, GTPases and ATPases. These assembly factors promote pre-rRNA folding and processing, remodeling of protein–protein and RNA–protein networks, nuclear export and quality control” [16].  While I suspect that some structural components of the ribosome and the nuclear pore may have functions as monomeric polypeptides, and so could be considered as proteins, at this point it is best (most accurate) to assume that they are polypeptides, components of proteins and larger, molecular machines (past post). 

We can, of course, continue to consider the roles of common folding motifs,  arising from the chemistry of the peptide bond and the environment within and around the assembling protein, in the context of protein structure [17, 18], The knottier problem is how to help students recognize how functional entities, proteins and molecular machines, together with the coupled reaction systems that drive them and the molecular interactions that regulate them, function. How mutations, alleleic variations, and various environmentally induced perturbations influence the behaviors of cells and organisms, and how they generate normal and pathogenic phenotypes. Such a view emphasizes the dynamics of the living state, and the complex flow of information out of DNA into networks of molecular machines and reaction systems. 

: Thanks to Michael Stowell for feedback and suggestions and Jon Van Blerkom for encouragement.  All remaining errors are mine. 


  1. Recently emerged from the labs of Martin Raff and Lee Rubin – Martin is one of the founding authors of the transformative “molecular biology of the cell” textbook. 
  2. Or rather quite over-simplistic, as it ignore complexities arising from differential splicing, alternative promoters, and genes encoding non-polypeptide encoding RNAs. 

Molecular machines and the place of physics in the biology curriculum

The other day, through no fault of my own, I found myself looking at the courses required by our molecular biology undergraduate degree program. I discovered a requirement for a 5 credit hour physics course, and a recommendation that this course be taken in the students’ senior year – a point in their studies when most have already completed their required biology courses.  Befuddlement struck me, what was the point of requiring an introductory physics course in the context of a molecular biology major?  Was this an example of time-travel (via wormholes or some other esoteric imagining) in which a physics course in the future impacts a students’ understanding of molecular biology in the past?  I was also struck by the possibility that requiring such a course in the students’ senior year would measurably impact their time to degree.  

In a search for clarity and possible enlightenment, I reflected back on my own experiences in an undergraduate biology degree program – as a practicing cell and molecular  biologist, I was somewhat confused. I could not put my finger on the purpose of our physics requirement, except perhaps the admirable goal of supporting physics graduate students. But then, after feverish reflections on the responsibilities of faculty in the design of the courses and curricula they prescribe for their students and the more general concepts of instructional (best) practice and malpractice, my mind calmed, perhaps because I was distracted by an article on Oxford Nanopore’s MinION (→) “portable real-time device for DNA and RNA sequencing”, a device that plugs into the USB port on one’s laptop! Distracted from the potentially quixotic problem of how to achieve effective educational reform at the undergraduate level, I found myself driven on by an insatiable curiosity (or a deep-seated insecurity) to ensure that I actually understood how this latest generation of DNA sequencers worked. This led me to a paper by Meni Wanunu (2012. Nanopores: A journey towards DNA sequencing)[1].  On reading the paper, I found myself returning to my original belief, yes, understanding physics is critical to developing a molecular-level understanding of how biological systems work, BUT it was just not the physics normally inflicted upon (required of) students [2]. Certainly this was no new idea.  Bruce Alberts had written on this topic a number of times, most dramatically in his 1989 paper “The cell as a collection of molecular machines” [3].  Rather sadly, and not withstanding much handwringing about the importance of expanding student interest in, and understanding of, STEM disciplines, not much of substance in this area has occurred. While (some minority of) physics courses may have adopted active engagement pedagogies, in the meaning of Hake [4], most insist on teaching macroscopic physics, rather than to focus on, or even to consider, the molecular level physics relevant to biological systems, explicitly the physics of protein machines in a cellular (biological) context. Why sadly? Because conventional, that is non-biologically relevant introductory physics and chemistry courses, all too often serve the role of a hazing ritual, driving many students out of the biological sciences [5], in part I suspect because they often seem irrelevant to students’ interests in the workings of biological systems.[6].  

Nanopore’s sequencer and Wanunu’s article got me thinking again about biological machines, of which there are a great number, ranging from pumps, propellers, and oars to  various types of transporters, molecular truckers that move chromosomes, membrane vesicles, and parts of cells with respect to one another, to DNA detanglers, protein unfolders, and molecular recyclers (→).  The Nanopore sequencer works because as a single strand of DNA (or RNA) moves through a narrow pore, the different bases (A,C,T,G) occlude the pore to different extents, allowing different numbers of ions, different amounts of current, to pass through the pore. These current differences can be detected, and allows for nucleotide sequence to be read as the nucleic acid strand moves through the pore. Understanding the process involves understanding how molecules move, that is the physics of molecular collisions and energy transfer, how proteins and membranes allow and restrict ion movement, and the impact of chemical gradients and electrical fields across a membrane on molecular movements  – all physical concepts of widespread significance in biological systems.  Such ideas can be extended to the more general questions of how molecules move within the cell, and the effects of molecular size and inter-molecular interactions within a concentrated solution of proteins, protein polymers, lipid membranes, and nucleic acids, such as described in Oliverira et al., Increased cytoplasm viscosity hampers aggregate polar segregation in Escherichia coli [7].  At the molecular level the processes, while biased by electric fields (potentials) and concentration gradients, are stochastic (noisy). Understanding of stochastic processes is difficult for students [8], but critical to developing an appreciation of how such processes can lead to phenotypic differences between cells with the same genotypes (previous post) and how such noisy processes are managed by the cell and within a multicellular organism.    

As path leads on to path, I found myself considering the (← adapted from Joshi et al., 2017) spear-chucking protein machine present in the pathogenic bacteria Vibrio cholerae; this molecular machine is used to inject toxins into neighbors that the bacterium happens to bump into (see Joshi et al., 2017. Rules of Engagement: The Type VI Secretion System in Vibrio cholerae)[9].  The system is complex and acts much like a spring-loaded and rather “inhumane” mouse trap.  This is one of a  number of bacterial type VI systems, and “has structural and functional homology to the T4 bacteriophage tail spike and tube” – the molecular machine that injects bacterial cells with the virus’s genetic material, its DNA.

Building the bacterium’s spear-based injection system is controlled by a social (quorum sensing) system (previous post). One of the ways that such organisms determine whether they are alone or living in an environment crowded with other organisms. During the process of assembly, potential energy, derived from various chemically coupled, thermodynamically favorable reactions, is stored in both type VI “spears” and the contractile (nucleic acid injecting) tails of the bacterial viruses (phage). Understanding the energetics of this process, for example, how coupling thermodynamically favorable chemical reactions, such as ATP hydrolysis, or physico-chemical reactions such as the diffusion of ions down an electrochemical gradient, can be used to set these “mouse traps”, and understandingwhere the energy goes when the traps are sprung is central to students’ understanding of these and a wide range of other molecular machines. 

Energy stored in such molecular machines during their assembly can be used to move the cell. As an example, another bacterial system generates contractile (type IV pili) filaments; the contraction of such a filament can allow “the bacterium to move 10 000 times its own body weight, which results in rapid movement” (see Berry & Belicic 2015. Exceptionally widespread nanomachines composed of type IV pilins: the prokaryotic Swiss Army knives)[10].  The contraction of such a filament has also been found to be used to import DNA into the cell, the first step in the process of horizontal gene transfer.  In other situations (other molecular machines) such protein filaments access thermodynamically favorable processes to rotate, acting like a propeller, driving cellular movement. 


During my biased random walk through the literature, I came across another, but molecularly distinct, machine used to import DNA into Vibrio (see Matthey & Blokesch 2016. The DNA-Uptake Process of Naturally Competent Vibrio cholerae)[11]. This molecular machine enables the bacterium to import DNA from the environment, released, perhaps, from a neighbor killed by its spear.  In this system (adapted from Matthey & Bioesch  et al., 2017 →), the double stranded DNA molecule is first transported through the bacterium’s outer membrane (“OM”); the DNA’s two strands are then separated, and one strand passes through a channel protein through the inner (plasma) membrane, and into the cytoplasm, where it can interact with the bacterium’s  genomic DNA.    

The value of introducing students to the idea of molecular machines is that it can be used to demystify how biological systems work, how such machines carry out specific functions, whether moving the cell or recognizing and repairing damaged DNA.  If physics matters in biological curriculum, it matters for this reason – it establishes th e core premise of biology that organisms are not driven by “vital” forces, but by prosaic physiochemical ones.  At the same time, the molecular mechanisms behind evolution, such as mutation, gene duplication, and genomic reorganization, provide the means by which new structures emerge from pre-existing ones, yet many is the molecular biology degree program that does not include an introduction to evolutionary mechanisms in its required course sequence – imagine that, requiring physics but not evolution?  [see 12].One final point regarding requiring students to take a biologically relevant physics course early in their degree program is that it can be used to reinforce what I think is a critical and often misunderstood point. While biological systems rely on molecular machines, we (and by we I mean all organisms) are NOT machines, no matter what physicists might postulate – see We Are All Machines That Think.  We are something different and distinct. Our behaviors and our feelings, whether ultimately understandable or not, emerge from the interaction of genetically encoded, stochastically driven non-equilibrium systems, modified through evolutionary, environmental, social, and a range of unpredictable events occurring in an uninterrupted, and basically undirected fashion for ~3.5 billion years.  While we are constrained, we are more, in some weird and probably ultimately incomprehensible way.

Making education matter in higher education

It may seem self-evident that providing an effective education, the type of educational experiences that lead to a useful bachelors degree and serve as the foundation for life-long learning and growth, should be a prime aspirational driver of Colleges and Universities (1).  We might even expect that various academic departments would compete with one another to excel in the quality and effectiveness of their educational outcomes; they certainly compete to enhance their research reputations, a competition that is, at least in part, responsible for the retention of faculty, even those who stray from an ethical path. Institutions compete to lure research stars away from one another, often offering substantial pay raises and research support (“Recruiting or academic poaching?”).  Yet, my own experience is that a department’s performance in undergraduate educational outcomes never figures when departments compete for institutional resources, such as supporting students, hiring new faculty, or obtaining necessary technical resources (2).

 I know of no example (and would be glad to hear of any) of a University hiring a professor based primarily on their effectiveness as an instructor (3).

In my last post, I suggested that increasing the emphasis on measures of departments’ educational effectiveness could help rebalance the importance of educational and research reputations, and perhaps incentivize institutions to be more consistent in enforcing ethical rules involving research malpractice and the abuse of students, both sexual and professional. Imagine if administrators (Deans and Provosts and such) were to withhold resources from departments that are performing below acceptable and competitive norms in terms of undergraduate educational outcomes?

Outsourced teaching: motives, means and impacts

Sadly, as it is, and particularly in many science departments, undergraduate educational outcomes have little if any impact on the perceived status of a department, as articulated by campus administrators. The result is that faculty are not incentivized to, and so rarely seriously consider the effectiveness of their department’s course requirements, a discussion that would of necessity include evaluating whether a course’s learning goals are coherent and realistic, whether the course is delivered effectively, whether it engages students (or is deemed irrelevant), and whether students’ achieve the desired learning outcomes, in terms of knowledge and skills achieved, including the ability to apply that knowledge effectively to new situations.  Departments, particularly research focussed (dependent) departments, often have faculty with low teaching loads, a situation that incentivizes the “outsourcing” of key aspects of their educational responsibilities.  Such outsourcing comes in two distinct forms, the first is requiring majors to take courses offered by other departments, even if such courses are not well designed, delivered, or (in the worst cases) relevant to the major.  A classic example is to require molecular biology students to take macroscopic physics or conventional calculus courses, without regard to whether the materials presented in these courses is ever used within the major or the discipline.  Expecting a student majoring in the life sciences to embrace a course that (often rightly) seems irrelevant to their discipline can alienate a student, and poses an unnecessary obstacle to student success, rather than providing students with needed knowledge and skills.  Generally, the incentives necessary to generate a relevant course, for example, a molecular level physics course that would engage molecular biology students, are simply not there.  A version of this situation is to require courses that are poorly designed or delivered (general chemistry is often used as the poster child for such a course). These are courses that have high failure rates, sometimes justified in terms of “necessary rigor” when in fact better course design could (and has) resulted in lower failure rates and improved learning outcomes.  In addition, there are perverse incentives associated with requiring “weed out” courses offered by other departments, as they reduce the number of courses a department’s faculty needs to teach, and can lead to fewer students proceeding into upper division courses.

The second type of outsourcing involves excusing tenure track faculty from teaching introductory courses, and having them replaced by lower paid instructors or lecturers.  Independently of whether instructors, lecturers, or tenure track professors make for better teaching, replacing faculty with instructors sends an implicit message to students.  At the same time, the freedom of instructors/lecturers to adopt an effective (socratic) approach to teaching is often severely constrained; common exams can force classes to move in lock step, independently of whether that pace is optimal for student engagement and learning. Generally, instructors/lecturers do not have the freedom to adjust what they teach, to modify the emphasis and time they spend on specific topics in response to their students’ needs. How an instructor instructs their students suffers when teachers do not have the freedom to customize their interactions with students in response to where they are intellectually.  This is particularly detrimental in the case of underrepresented or underprepared students. Generally, a flexible and adaptive approach to instruction (including ancillary classes on how to cope with college: see An alternative to remedial college classes gets results) can address many issues, and bring the majority of students to a level of competence, whereas tracking students into remedial classes can succeed in driving them out of a major or college (see Colleges Reinvent Classes to Keep More Students in Science and Redesigning a Large-Enrollment Introductory Biology Course and Does Remediation Work for All Students? )

How to address this imbalance, how can we reset the pecking order so that effective educational efforts actually matter to a department? 

My (modest) suggestion is to base departmental rewards on objective measures of educational effectiveness.   And by rewards I mean both at the level of individuals (salary and status) as well as support for graduate students, faculty positions, start up funds, etc.  What if, for example, faculty in departments that excel at educating their students received a teaching bonus, or if the number of graduate students within a department supported by the institution was determined not by the number of classes these graduate students taught (courses that might not be particularly effective or engaging) but rather by a departments’ undergraduate educational effectiveness, as measured by retention, time to degree, and learning outcomes (see below)?  The result could well be a drive within a department to improve course and curricular effectiveness to maximize education-linked rewards.  Given that laboratory courses, the courses most often taught by science graduate students, are multi-hour schedule disrupting events, of limited demonstrable educational effectiveness, that complicate student course scheduling, removing requirements for lab courses deemed unnecessary (or generating more effective versions), would be actively rewarded (of course, sanctions for continuing to offer ineffective courses would also be useful, but politically more problematic.)

A similar situation applies when a biology department requires its majors to take 5 credit hour physics or chemistry courses.  Currently it is “easy” for a department to require its students to take such courses without critically evaluating whether they are “worth it”, educationally.  Imagine how a department’s choices of required courses would change if the impact of high failure rates (which I would argue is a proxy for poorly designed  and delivered courses) directly impacted the rewards reaped by a department. There would be an incentive to look critically at such courses, to determine whether they are necessary and if so, well designed and delivered. Departments would serve their own interests if they invested in the development of courses  that better served their disciplinary goals, courses likely to engage their students’ interests.

So how do we measure a department’s educational efficacy?

There are three obvious metrics: i) retention of students as majors (or in the case of “service courses” for non-majors, whether students master what it is the course claims to teach); ii) time to degree (and by that I mean the percentage of students who graduate in 4 years, rather than the 6 year time point reported in response to federal regulations (six year graduation rate | background on graduation rates); and iii) objective measures of student learning outcomes attained and skills achieved. The first two are easy, Universities already know these numbers.  Moreover they are directly influenced by degree requirements – requiring students to take boring and/or apparently irrelevant courses serves to drive a subset of students out of a major.  By making courses relevant and engaging, more students can be retained in a degree program. At the same time, thoughtful course design can help students  pass through even the most rigorous (difficult) of such courses. The third, learning outcomes, is significantly more challenging to measure, since universal metrics are (largely) missing or superficial.  A few disciplines, such as chemistry, support standardized assessments, although one could argue with what such assessments measure.  Nevertheless, meaningful outcomes measures are necessary, in much the same way that Law and Medical boards and the Fundamentals of Engineering exam serve to help insure (although they do not guarantee) the competence of practitioners. One could imagine using parts of standardized exams, such as discipline specific GRE exams, to generate outcomes metrics, although more informative assessment instruments would clearly be preferable. The initiative in this area could be taken by professional societies, college consortia (such as the AAU), and research foundations, as a critical driver for education reform, increased effectiveness, and improved cost-benefit outcomes, something that could help address the growing income inequality in our country and make success in higher education an important factor contributing to an institution’s reputation.


A footnote or two…
1. My comments are primarily focused on research universities, since that is where my experience lies; these are, of course, the majority of the largest universities (in a student population sense).
2. Although my experience is limited, having spent my professorial career at a single institution, conversations with others leads me to conclude that it is not unique.
3. The one obvious exception would be the hiring of  coaches of sports teams, since their success in teaching (coaching) is more directly discernible and impactful on institutional finances and reputation).
minor edits – 16 March 2020

Reverse Dunning-Kruger effects and science education

The Dunning-Kruger (DK) effect is the well-established phenomenon that people tend to over estimate their understanding of a particular topic or their skill at a particular task, often to a dramatic degree [link][link]. We see examples of the DK effect throughout society; the current administration (unfortunately) and the nutritional supplements / homeopathy section of Whole Foods spring to mind as examples. But there is a less well-recognized “reverse DK” effect, namely the tendency of instructors, and a range of other public communicators, to over-estimate what the people they are talking to are prepared to understand, appreciate, and accurately apply. The efforts of science communicators and instructors can be entertaining but the failure to recognize and address the reverse DK effect results in ineffective educational efforts. These efforts can themselves help generate the illusion of understanding in students and the broader public (discussed here). While a confused understanding of the intricacies of cosmology or particle physics can be relatively harmless in their social and personal implications, similar misunderstandings become personally and publicly significant when topics such as vaccination, alternative medical treatments, and climate change are in play.

There are two synergistic aspects to the reverse DK effect that directly impact science instruction: the need to understand what one’s audience does not understand together with the need to clearly articulate the conceptual underpinnings needed to understand the subject to be taught. This is in part because modern science has, at its core, become increasingly counter-intuitive over the last approximately 100 years or so, a situation that can cause serious confusions that educators must address directly and explicitly. The first reverse DK effect involves the extent to which the instructor (and by implication the course and textbook designer) has an accurate appreciation of what students think or think they know, what ideas they have previously been exposed to, and what they actually understand about the implications of those ideas.  Are they prepared to learn a subject or does the instructor first have to acknowledge and address conceptual confusions and build or rebuild base concepts?  While the best way to discover what students think is arguably a Socratic discussion, this only rarely occurs for a range of practical reasons. In its place, a number of concept inventory-type testing instruments have been generated to reveal whether various pre-identified common confusions exist in students’ thinking. Knowing the results of such assessments BEFORE instruction can help customize how the instructor structures the learning environment and content to be presented and whether the instructor gives students the space to work with these ideas to develop a more accurate and nuanced understanding of a topic.  Of course, this implies that instructors have the flexibility to adjust the pace and focus of their classroom activities. Do they take the time needed to address student issues or do they feel pressured to plow through the prescribed course content, come hell, high water, or cascading student befuddlement.

A complementary aspect of the reverse DK effect, well-illustrated in the “why magnets attract” interview with the physicist Richard Feynman, is that the instructor, course designer, or textbook author(s) needs to have a deep and accurate appreciation of the underlying core knowledge necessary to understand the topic they are teaching. Such a robust conceptual understanding makes it possible to convey the complexities involved in a particular process and explicitly values appreciating a topic rather than memorizing it.  It focuses on the general, rather than the idiosyncratic. A classic example from many an introductory biology course is the difference between expecting students to remember the steps in glycolysis or the Krebs cycle reaction system, as opposed to the general principles that underlie the non-equilibrium reaction networks involved in all biological functions, a reaction network based on coupled chemical reactions and governed by the behaviors of thermodynamically favorable and unfavorable reactions. Without a explicit discussion of these topics, all too often students are required to memorize names without understanding the underlying rationale driving the processes involved; that is, why the system behaves as it does.  Instructors also give false “rubber band” analogies or heuristics to explain complex phenomena (see Feynman video 6:18 minutes in). A similar situation occurs when considering how molecules come to associate and dissociate from one another, for example in the process of regulating gene expression or repairing mutations in DNA. Most textbooks simply do not discuss the physiochemical processes involved in binding specificity, association, and dissociation rates, such as the energy changes associated with molecular interactions and thermal collisions (don’t believe me? look for yourself!). But these factors are essential for a student to understand the dynamics of gene expression [link], as well as the specificity of modern methods involved in genetic engineering, such as restriction enzymes, polymerase chain reaction, and CRISPR CAS9-mediated mutagenesis. By focusing on the underlying processes involved we can avoid their trivialization and enable students to apply basic principles to a broad range of situations. We can understand exactly why CRISPR CAS9-directed mutagenesis can be targeted to a single site within a multibillion-base pair genome.

Of course, as in the case of recognizing and responding to student misunderstandings and knowledge gaps, a thoughtful consideration of underlying processes takes course time, time that trades the development of a working understanding of core processes and principles for broader “coverage” of frequently disconnected facts, the memorization and regurgitation of which has been privileged over understanding why those facts are worth knowing. If our goal is for students to emerge from a course with an accurate understanding of the basic processes involved rather than a superficial familiarity with a plethora of unrelated facts, however, a Socratic interaction with the topic is essential. What assumptions are being made, where do they come from, how do they constrain the system, and what are their implications?  Do we understand why the system behaves the way it does? In this light, it is a serious educational mystery that many molecular biology / biochemistry curricula fail to introduce students to the range of selective and non-selective evolutionary mechanisms (including social and sexual selection – see link), that is, the processes that have shaped modern organisms.

Both aspects of the reverse DK effect impact educational outcomes. Overcoming the reverse DK effect depends on educational institutions committing to effective and engaging course design, measured in terms of retention, time to degree, and a robust inquiry into actual student learning. Such an institutional dedication to effective course design and delivery is necessary to empower instructors and course designers. These individuals bring a deep understanding of the topics taught and their conceptual foundations and historic development to their students AND must have the flexibility and authority to alter the pace (and design) of a course or a curriculum when they discover that their students lack the pre-existing expertise necessary for learning or that the course materials (textbooks) do not present or emphasize necessary ideas. Unfortunately, all too often instructors, particularly in introductory level college science courses, are not the masters of their ships; that is, they are not rewarded for generating more effective course materials. An emphasis on course “coverage” over learning, whether through peer-pressure, institutional apathy, or both, generates unnecessary obstacles to both student engagement and content mastery.  To reverse the effects of the reverse DK effect, we need to encourage instructors, course designers, and departments to see the presentation of core disciplinary observations and concepts as the intellectually challenging and valuable endeavor that it is. In its absence, there are serious (and growing) pressures to trivialize or obscure the educational experience – leading to the socially- and personally-damaging growth of fake knowledge.