← writing

Metacognition

My mid-terms just ended and over the last few days I have been particularly intrigued by The Mentalist. Seeing Patrick Jane reading people made me much more curious about body language and microfacial expressions. So instead of finishing my supplements, I spent (wasted) 3 days binge watching these episodes. That was a sad affair. I have finished off the series (Thankfully).

Finishing it off got me thinking about this deep desire of mine to keep binge-watching for the rest of my life. Of course, it is highly impractical and not at all a part of my goals or interest, so I needed to get rid of it. That’s when I came across the basic dopamine based theories. I had it under control in 5 days. I started studying again.

There was a caveat though.

I was not able to learn effectively.

Now, that was an interesting problem. I had solved the problem of starting. I could sit down, put the distractions away and study. But somewhere between sitting down with a book and actually knowing the material, something was going wrong.

So I went down another rabbithole to figure out why.

That is when I came across metacognition.

The problem with knowing that you know

Metacognition, in its simplest form, is thinking about your own thinking. It is the ability to monitor what you know, what you do not know, and how effectively you are approaching a problem. In learning, it becomes particularly important because studying is not just about acquiring information; it is also about deciding whether that information has actually been acquired.

This sounds trivial until you realise how badly we can estimate our own learning.

Suppose I study a chapter for two hours. At the end of it, I have read everything, highlighted the important parts, understood the explanations while looking at them, and generally feel like I have done a good job.

Let me call my subjective estimate of how much I have learned JJ, for judgment of learning.

The actual amount I can demonstrate later is PP, for performance.

Ideally,

J≈PJ \approx P

But there is no reason for that to automatically be true.

If I think I have learned 90%90\% of the material but can only correctly reproduce 60%60\% of it later, then my problem is not simply that I have forgotten things. My problem is that I didn't know that I didn't know them.

Mathematically, we can define a very simple metacognitive error as

E=J−PE = J-P

where EE is the difference between what I think I know and what I can actually demonstrate.

If

E>0E>0

I am overestimating myself.

If

E<0E<0

I am underestimating myself.

And if

E≈0E\approx 0

then my judgment is reasonably calibrated.

This is basically what researchers mean when they talk about calibration in metacognitive monitoring: how closely a person's confidence or judgment corresponds to their actual performance.

The interesting part is that this error is not just an interesting number. It affects what I do next.

If I believe I know something, I stop studying it.

If I believe I don't know something, I study it again.

So my estimate of my own knowledge controls the allocation of my time.

We can represent this rather simply:

Judgment→Study Decision→Performance\text{Judgment} \rightarrow \text{Study Decision} \rightarrow \text{Performance}

The problem is that if the judgment is wrong, the entire chain can go wrong.

Imagine I have ten topics to study. I incorrectly believe that Topics 1,2,3,41,2,3,4 are already mastered and that Topics 5,6,7,8,9,105,6,7,8,9,10 need work. I spend all my remaining time on 55 through 1010.

But suppose reality looks like this:

P=[0.95,0.90,0.55,0.50,0.70,0.65,0.80,0.60,0.75,0.85]P = [0.95,0.90,0.55,0.50,0.70,0.65,0.80,0.60,0.75,0.85]

and my subjective judgment was:

J=[0.95,0.95,0.90,0.90,0.70,0.65,0.80,0.60,0.75,0.85]J = [0.95,0.95,0.90,0.90,0.70,0.65,0.80,0.60,0.75,0.85]

I have completely missed Topics 33 and 44.

They felt familiar.

And familiarity is a dangerous thing.

Familiarity is not mastery

This, I think, was a large part of my problem.

When I read something repeatedly, the information becomes easier to process. The terminology becomes familiar. I know where the important paragraph is. I recognise the equation. When I see the solution, it makes sense.

And then I think:

"Yeah, I know this."

But do I?

There is a fundamental difference between recognition and recall.

When the textbook is open, I am receiving the information.

When the textbook is closed, I have to generate it.

That difference is exactly why retrieval practice is so useful. Instead of asking myself whether something looks familiar, I can ask whether I can actually produce the answer without being given the answer first.

This gives me a much better measurement of PP.

The process becomes something like:

Study→Retrieve→Measure→Update\boxed{ \text{Study} \rightarrow \text{Retrieve} \rightarrow \text{Measure} \rightarrow \text{Update} }

And suddenly studying stops being this vague activity where I sit at a desk for an arbitrary number of hours.

It becomes a feedback system.

Learning as a feedback system

This is probably the part that interested me the most.

If I think about learning mathematically, there is an initial state:

KtK_t

where KtK_t represents my knowledge at time tt.

I use some learning strategy StS_t, spend some amount of effort utu_t, and after some time I get a new state:

Kt+1=Kt+ΔKtK_{t+1}=K_t+\Delta K_t

The problem is that I don't directly observe KtK_t.

I only observe imperfect signals of it.

That is where metacognition comes in.

I make an estimate:

K^t\hat{K}_t

where the hat means my estimate of my actual knowledge.

The important quantity is therefore:

et=Kt−K^te_t = K_t-\hat{K}_t

The problem is that KtK_t is hidden. I don't know my actual state until I test it.

So I need measurements.

A question.

A problem.

A blank sheet of paper.

An explanation from memory.

A practice test.

Anything that forces the information out rather than simply putting it in front of me.

Then I get something much closer to an observation of KtK_t, and I can update my estimate.

This is basically the logic behind metacognitive monitoring: I am constantly trying to estimate an internal state that I cannot observe directly.

And that is a surprisingly interesting problem.

The error signal

There is another connection that I found particularly interesting because of my original dopamine rabbit hole.

In reinforcement learning, one of the classic ideas is the prediction error. Very roughly, the error can be represented as

δt=rt−r^t\delta_t = r_t-\hat{r}_t

where rtr_t is the reward that actually occurs and r^t\hat{r}_t is the reward that was expected.

The difference between the two is the prediction error.

Dopamine signalling has been extensively studied in relation to reward-prediction errors and learning, although the modern picture is considerably more complicated than the popular "dopamine = pleasure" explanation. Contemporary research suggests dopamine signals carry information relevant to learning, motivation, uncertainty and prediction, rather than functioning as a simple pleasure meter.

And I started wondering whether there is a similar structure in learning.

I predict:

P^=0.90\hat{P}=0.90

I expect that after studying a chapter I will answer 90%90\% of the questions correctly.

Then I actually test myself and get:

P=0.60P=0.60

So my learning prediction error is

ϵ=P−P^\epsilon = P-\hat{P}

and therefore

ϵ=0.60−0.90=−0.30\epsilon = 0.60-0.90=-0.30

That −0.30-0.30 is useful.

It tells me that my internal model of my own learning was wrong.

And that is where metacognition becomes more than simply "thinking about thinking." The purpose of monitoring is ultimately to control what happens next. If I discover that I have overestimated my understanding, I can change the amount of time I allocate to that topic, change my study method, or test myself again.

In fact, you could imagine a very simple update rule:

P^t+1=P^t+α(Pt−P^t)\hat{P}_{t+1} = \hat{P}_t+\alpha(P_t-\hat{P}_t)

where 0<α≤10<\alpha\leq 1 represents how strongly I update my belief based on new evidence.

If I consistently predict 90%90\% and repeatedly score 60%60\%, my estimate should eventually move towards reality.

That, to me, is the interesting part.

Metacognition is essentially the attempt to improve the model you have of your own mind.

But why was I studying badly?

Looking back, I think I had solved only half of my original problem.

The binge-watching problem was primarily a problem of behaviour and attention. I had to get myself back into a state where studying was possible.

But once I had done that, another problem appeared.

I had no good feedback mechanism.

I was measuring the wrong variable.

I was measuring:

T=time spent studyingT = \text{time spent studying}

when what I actually cared about was something closer to:

L=learning achievedL = \text{learning achieved}

And these are obviously not the same thing.

It is entirely possible that

T1>T2T_1>T_2

while

L1<L2L_1<L_2

In other words, I can study for longer and learn less.

This seems painfully obvious when written mathematically, but I think most of us behave as though

L∝TL\propto T

as if learning were simply proportional to the number of hours spent at a desk.

It isn't.

The relationship depends on the strategy, the material, prior knowledge, attention, sleep, motivation, retrieval, feedback and a ridiculous number of other variables.

So perhaps the better model is something like

L=f(T,S,R,F,A,…)L=f(T,S,R,F,A,\ldots)

where SS is the study strategy, RR is retrieval, FF is feedback, and AA represents attention.

The exact function is obviously not something I can write down.

But the point is that time is only one variable.

And I had been optimising the easiest variable to measure.

The strange part about being wrong

There is something slightly uncomfortable about metacognition because it forces you to accept that your own subjective experience is not necessarily a reliable measurement.

I can feel like I understand something and still not understand it.

I can feel like I had a productive day and have very little to show for it.

I can feel like a topic is easy because I have seen it many times.

And I can also feel like something is difficult when I am actually learning it quite well.

That last one is particularly interesting.

Some effective learning techniques create desirable difficulty. Retrieval practice, spacing and interleaving can make learning feel harder in the short term while improving later performance. So the subjective sensation of difficulty is not a sufficient measure of whether learning is working.

This means I need to separate two things:

How learning feels\text{How learning feels}

from

How learning performs\text{How learning performs}

The first is subjective.

The second can be measured.

And when the two disagree, I should probably trust the measurement.

So what actually changes?

I don't think the conclusion is that I should now create some ridiculously complicated productivity system with twelve spreadsheets and a Bayesian model of my brain.

That would probably become another form of procrastination.

The practical change is much simpler.

Before studying, I should know what I am trying to learn.

During studying, I should occasionally ask myself whether I actually understand it.

After studying, I should test myself without the material in front of me.

Then I should compare what I thought I knew with what I could actually demonstrate.

In the simplest possible form:

Predict→Study→Retrieve→Measure→Update\boxed{ \text{Predict} \rightarrow \text{Study} \rightarrow \text{Retrieve} \rightarrow \text{Measure} \rightarrow \text{Update} }

And then repeat.

That is a much better learning loop than simply:

Study→Study→Study\text{Study}\rightarrow\text{Study}\rightarrow\text{Study}

because the second system has no mechanism for telling me whether it is working.

Back to Patrick Jane

It is slightly ironic that all of this started with The Mentalist.

I became interested in Patrick Jane because of his ability to observe other people. I wanted to understand body language, facial expressions and all those tiny signals that supposedly reveal what someone is thinking.

Then I ended up learning that there is another person whose signals I should probably be paying much more attention to:

myself.

Why did I keep watching another episode when I knew I had work to do?

Why did I think I was learning when I was mostly recognising?

Why did I measure productivity in hours rather than outcomes?

Why did I stop studying a topic because it felt familiar?

Why did I not test the thing I was supposedly learning?

These are all questions about the same underlying problem: how accurately can I observe my own internal state?

And perhaps that is what makes metacognition so interesting to me.

I started this whole rabbit hole because I wanted to understand why I could not study effectively even after I had managed to stop wasting my time binge-watching.

I thought the problem was motivation.

Then I thought it was dopamine.

Then I thought it was discipline.

It turns out there was another layer underneath all of that.

I didn't just need to control my behaviour.

I needed to know whether my behaviour was actually producing the result I wanted.

And that requires feedback.

Maybe learning, at least in part, is just this:

Prediction→Action→Error→Correction\boxed{ \text{Prediction} \rightarrow \text{Action} \rightarrow \text{Error} \rightarrow \text{Correction} }

The better the feedback, the better the correction.

And the better the correction, the closer the prediction gets to reality.

So perhaps the real skill isn't just learning.

It is learning how badly you are learning, noticing it early, and having enough metacognition to change course.

if you have an argument, a disagreement, or something worth discussing, reach out at arnavd371[at]gmail[dot]com or arnav[at]aethra[dot]co[dot]in.