Edit History (Oldest to Newest)
Version: 1
Fields Changed (Original)
Updated
Content

"Imagine that we've got two separate research groups both testing the hypothesis that all-10s solve 2-4-6 with 50% propensity, or alternatively, less than 50% propensity.  They each don't know the other group exists; however, they both use the - hypothetically for this thought experiment - universally standard rule that 'less than 50% propensity' is of course best-modeled in practice by three equally weighted 'probability point-masses' on 0.1, 0.2, and 0.4."

"The first group reports a likelihood of 0.02 on the less-than-50% metahypothesis, and a likelihood of 0.031 on the 50% hypothesis."

"The second group reports the same thing."

"The way we've set up the hypotheses being reported on, we cannot just multiply the two likelihoods together.  The task of combining evidence from different 'published-experimental-reports' is now a big complicated deal requiring us to recheck their original data and redo all their calculations."

"Alternatively, if they had both reported likelihoods of 0.007, 0.02, 0.035, and 0.031, on the distinct hypotheses of 10%, 20%, 40%, and 50% propensity respectively, we could have just multiplied the likelihoods together from both groups, and our ability to accumulate data from across multiple experiments would be vastly simplified."

"Of which it is said out of dath ilan, to those dath ilani children who need to hear it:  Different 'effect-sizes' are different hypotheses."

"That, Carissa, Pilar, is why we can't just have the hypothesis that all-14s have at least five times the propensity of all-10s to solve 2-4-6 in 30 minutes.  We can look at the data and see if that actually happened or not.  But as soon as we try to figure out the exact likelihood that it happened, we are cast into a nightmarish multiverse of different ways the world could be, such that the statement 'all-14s are more than five times as likely to solve in thirty as all-10s' is true about worlds like that, all of which have different likelihoods of yielding the data we saw."

"Like, just on this breakdown, that could be because the chances were .2 and 1.0, or .1 and .5, or .1 and .6, or .1 and .8.  And every one of those hypothetical propensity '2-tuples' will yield a different likelihood for the data we saw.  So you'd have to put a prior on their relative odds inside that metahypothesis bucket, before you could calculate the likelihood for the whole bucket."

"And then, actually seeing any data, would update the odds inside that bucket, which would change the likelihood for any future experiments, even if the replicators saw exactly the same data you did."

"Of which it is said, again:  Let different effect sizes be different hypotheses."

Version: 2
Fields Changed Content
Updated
Content

"Imagine that we've got two separate research groups both testing the hypothesis that all-10s solve 2-4-6 with 50% propensity, or alternatively, less than 50% propensity.  They each don't know the other group exists; however, they both use the - hypothetically for this thought experiment - universally standard rule that 'less than 50% propensity' is of course best-modeled in practice by three equally weighted 'probability point-masses' on 0.1, 0.2, and 0.4."

"The first group reports a likelihood of 0.02 on the less-than-50% metahypothesis, and a likelihood of 0.031 on the 50% hypothesis."

"The second group reports the same thing."

"The way we've set up the hypotheses being reported on, we cannot just multiply the two likelihoods together.  The task of combining evidence from different 'published-experimental-reports' is now a big complicated deal requiring us to recheck their original data and redo all their calculations."

"Alternatively, if they had both reported likelihoods of 0.007, 0.02, 0.035, and 0.031, on the distinct hypotheses of 10%, 20%, 40%, and 50% propensity respectively, we could have just multiplied the likelihoods together from both groups, and our ability to accumulate data from across multiple experiments would be vastly simplified."

"Of which it is said out of dath ilan, to those dath ilani children who need to hear it:  Different 'effect-sizes' are different hypotheses."

"That, Carissa, Pilar, is why we can't just have the hypothesis that all-14s have at least five times the propensity of all-10s to solve 2-4-6 in 30 minutes.  We can look at the data and see if that actually happened or not.  But as soon as we try to figure out the exact likelihood that it happened, we are cast into a nightmarish multiverse of different ways the world could be, such that the statement 'all-14s are more than five times as likely to solve in thirty as all-10s' is true about worlds like that, all of which have different likelihoods of yielding the data we saw."

"Like, just on this breakdown, that could be because the chances were .2 and 1.0, or .1 and .5, or .1 and .6, or .1 and .8.  And every one of those hypothetical propensity '2-tuples' will yield a different likelihood for the data we saw.  So you'd have to put a prior on their relative odds inside that metahypothesis bucket, before you could calculate the likelihood for the whole bucket."

"And then, actually seeing any data, would update the odds inside that bucket, which would change the likelihood for any future experiments, even if the replicators saw exactly the same data you did."

Version: 3
Fields Changed Content
Updated
Content

"Imagine that we've got two separate research groups both testing the hypothesis that all-10s solve 2-4-6 with 50% propensity, or alternatively, less than 50% propensity.  They each don't know the other group exists; however, they both use the - hypothetically for this thought experiment - universally standard rule that 'less than 50% propensity' is of course best-modeled in practice by three equally weighted 'probability point-masses' on 0.1, 0.2, and 0.4."

"The first group reports a likelihood of 0.02 on the less-than-50% metahypothesis, and a likelihood of 0.031 on the 50% hypothesis."

"The second group reports the same thing."

"The way we've set up the hypotheses being reported on, we cannot just multiply the two likelihoods together.  The task of combining evidence from different 'published-experimental-reports' is now a big complicated deal requiring us to recheck their original data and redo all their calculations."

"Alternatively, if they had both reported likelihoods of 0.007, 0.02, 0.035, and 0.031, on the distinct hypotheses of 10%, 20%, 40%, and 50% propensity respectively, we could have just multiplied the likelihoods together from both groups, and our ability to accumulate data from across multiple experiments would be vastly simplified."

"Of which it is said out of dath ilan, to those dath ilani children who need to hear it:  Different 'effect-sizes' are different hypotheses."

"That, Carissa, Pilar, is why we can't just have the hypothesis that all-14s have at least five times the propensity of all-10s to solve 2-4-6 in 30 minutes.  We can look at the data and see if that actually happened or not.  But as soon as we try to figure out the exact likelihood that it happened, we are cast into a nightmarish multiverse of different ways the world could be, such that the statement 'all-14s are more than five times as likely to solve in thirty as all-10s' is true about worlds like that, all of which have different likelihoods of yielding the data we saw."

"Like, just on this breakdown, that could be because the chances were .2 and 1.0, or .1 and .5, or .1 and .6, or .1 and .8.  And every one of those hypothetical propensity '2-tuples' will yield a different likelihood for whatever data we saw.  So you'd have to put a prior on their relative odds inside that metahypothesis bucket, before you could calculate the likelihood for the whole bucket."

"And then, actually seeing any data, would update the odds inside that bucket, which would change the likelihood for any future experiments, even if the replicators saw exactly the same data you did."