Monday, July 1, 2024

The Wisdom of Scientific Inquiry on Education

The Wisdom of Scientific Inquiry on Education

Gene V Glass
University of Colorado

This paper was written in 1970 and presented at a Convention of the National Association for Research on Science Teaching in Silver Spring, Maryland on 24 March 1971.

INTRODUCTION

A few years ago, Ralph Tyler assessed the state of research on science teaching and found it in need of a more theoretical, scientific perspective. My own assessment of the field now and then is diametrically opposed to Tyler's and, incidentally, to the position taken by Pella (1966) at about the same time Tyler wrote on this subject. For fear that a paraphrasing of Tyler's position might distort his meaning, I shall quote him at some length. Tyler's prescription for improved research in science education was as follows:

Theory is the all embracing end of basic research in seeking to provide a comprehensive map of the terrain of science education. Concepts are the smaller areas which comprise the total map, or to put the metaphor in another way, the complex of science education can be understood more readily by considering the concepts as major parts of the whole, and studying these parts in greater detail than is possible with the total.... The concepts and the dynamic models furnish the map which we seek in order to understand the factors involved in, and the process of science education. They form the major part of the theory....
My criticism of current research is its failure to be guided by, or to produce an adequate map of the factors and processes in science education.... The outlining, elaboration, and testing of such a map seems to me to be the necessary focus of our attention if we are to improve research in this field. (Tyler, 1967-1968, p. 43.)
Tyler went on to name eleven prominent areas of "the map": 1) the objectives of science education; 2) the teaching-learning process; 3) the organization of learning experiences; 4) the outcomes of science education; 5) the student's development; 6) the development of teachers; 7) the objectives of education for science teachers; 8) the teaching-learning process of teacher education; 9) the outcomes of teacher education; 10) the organization of the teacher's learning experiences; 11) the process of change in science education. In my opinion, we should not strive to make research on science education or education generally more scientific. Indeed, we who call ourselves educational researchers should turn away from elucidatory inquiry in all areas of education. This type of inquiry, directed toward the construction of theories or models for the understanding and explanation of phenomena, should be left to the social and natural sciences because it is currently unproductive in education and is a profligate expenditure of precious resources of time, money, and talent. We should turn instead toward evaluative inquiry of educational developments which are the creations of masterful teachers inspired by what reliable knowledge exists in psychology, sociology, and the other sciences concerning the educating process.

My argument depends heavily on a distinction between two types of inquiry -- elucidatory and evaluative. The contrast between elucidation and evaluation is well established in aesthetics. Elucidatory aesthetics is the attempt to explain what constitutes art, or more generally beauty. Evaluative aesthetics concerns the discrimination between good and bad art -- what is beautiful and what is less beautiful -- without explanation of why an art work is good or bad. It is necessary to turn attention to the definitional problem and the related problem of distinguishing between elucidatory and evaluative inquiry.

ELUCIDATORY AND EVALUATIVE INQUIRY

Elucidatory inquiry is the process of obtaining generalizable knowledge by contriving and testing claims about relationships among variables or generalizable phenomena. Apologies to mathematics, history, philosophy, etc. for the obvious empirical social science bias in the definition and the thinking to follow. This knowledge results in functional or statistical relationships, models, and ultimately theories. When the results of elucidatory inquiry are combined with knowledge of particular circumstances, one obtains explanations, in the sense of Braithwaite (1953).

Evaluative inquiry is the determination of the worth of a thing. In education, it involves obtaining information to judge the worth of a program, product, or procedure. According to Scriven (1967, p. 40), "The activity consists simply in gathering and combining of performance data with a weighted set of goal scales to yield either comparative or numerical ratings; and in the justification of (a) the data-gathering instruments,(b) the weightings, and (c) the selection of goals."

Elucidatory and evaluative inquiry have many defining characteristics. Each is only imperfectly correlated with the tendency of informed persons to call activity A "elucidatory" and activity B "evaluative" just as clinical psychologists use "anxiety" as a construct to differentiate instances of behavior in a way that is not perfectly reproduced by a single measure or defining characteristic. The conceptualizations of elucidatory and evaluative inquiry are enriched by the identification of any characteristic of inquiry which has a non-zero correlation with the tendency of intelligent persons to speak of "elucidation" or "evaluation" when discussing a particular inquiry activity. Scriven (1958, p. 175) referred to such terms as "cluster concepts" or "correlational concepts." Such concepts, e.g., "schizophrenia," are known by their indicators all of which are imperfectly related to them.

CHARACTERISTICS OF INQUIRY

Nine characteristics of inquiry which distinguish elucidation from evaluation are recognizable.

[The next eighteen paragraphs are based on a segment of the author's paper with Blaine R. Worthen entitled "Educational Inquiry and the Practice of Education" which will appear as a chapter in Shalock, H. Del (Ed.) Frameworks for Viewing Educational Research & Development, Diffusion and Evaluation. Monmouth, Oregon; Teaching Research Division of Oregon State System of Higher Education.(In press)]

1. Motivation of the Inquirer
Elucidatory and evaluative inquiry appear generally to be undertaken for different reasons. The former is pursued largely to satisfy curiosity; the latter is done to contribute to the solution of a practical problem. The theory builder is intrigued; evaluators (or at least their clients) are concerned with worth and value. Although the elucidatory inquirers may believe that their work has greater long-range payoff than the evaluator's, they are no less motivated by curiosity when performing their unique function. One must be nimble to avoid becoming bogged down in the seeming paradox that the policy decision to support basic inquiry because of its ultimate practical payoff does not imply that basic inquirers are pursuing practical ends in their daily work. Scriven (1969) argues that as regards research in mathematics, practical pay-off is increased to the extent that the mathematician is convinced he or she is not seeking practically significant results.

2. The Objective of the Search
Elucidatory and evaluative inquiry seek different ends. The former seeks conclusions; evaluation leads to decisions (see Tukey, 1960). Cronbach and Suppes (1969, pp. 20-21) distinguished between decision-oriented and conclusion-oriented inquiry.

In a decision-oriented study the investigator is asked to provide information wanted by a decision-maker: a school administrator, a government policy-maker, the manager of a project to develop a new biology textbook, or the like. The decision-oriented study is a commissioned study. Decision-makers believe that they need information to guide their actions and they pose the question to the investigator. The conclusion-oriented study, on the other hand, takes its direction from the investigator's commitments and hunches. The educational decision-maker can, at most, arouse the investigator's interest in a problem. The latter formulates the question, usually a general one rather than a question about a particular institution. The aim is to conceptualize and understand the chosen phenomenon; a particular finding is only a means to that end.

Conclusion-oriented inquiry is much like what is here referred to as elucidatory inquiry; decision-oriented inquiry typifies evaluative inquiry as well as any three words can.

3. Laws vs. Description
Closely related to the distinction between conclusion-oriented and decision-oriented are the familiar concepts of nomothetic (law-giving) and idiographic (descriptive of the particular). Elucidatory inquiry is the search for laws, i.e., statements of relationship among two or more variables or phenomena. Evaluation merely seeks to describe a particular thing with respect to one or more scales of value.

4. The Role of Explanation
Scientific explanations require scientific laws, and the disciplines related to education appear to far from discovery of the general laws on which explanations of incidents of schooling can be based. Explanations are not the goal of evaluation. A fully proper and useful evaluation can be conducted without producing an explanation of why the product or program being evaluated is good or bad or how it operates to produce its effects.

Elucidatory inquiry is characterized by a succession of studies in which greater control ("control" in the sense of the ability to manipulate specific components of independent variables) is exercised at each stage so that relationships among variables can be determined at more fundamental levels. Science seems to be an endless search for subsurface explanations, i.e., accountings of surface phenomena in terms of relationships among variables at a more subtle, covert level. Subsurface explanations are not sought after for themselves, or else the ultimate goal of science is some personal aberration, such as the science of Lawsonomy, a bizarre construction of an oddball inventor from Milwaukee. They are sought because greater precision, greater trust, the answer to the next "why?" always seems to be one stratum below the phenomena we see now. When ethologists find that the swallows return to Capistrano each March 19 because they are following the insects they feed on, they immediately ask why the insects return on March 19; and so it goes. The history of learning psychology is an excellent illustration of the continual search for subsurface explanations, though countless examples could be found in any science. Psychologists have sought increasingly more fundamental explanations of learning along a path that has led through the Law ofEffect, unravelings of the nature of reward, drive states, secondary reinforcement, and which leads inexorably toward brain physiology and the void beyond. Elucidatory inquiry chases subsurface explanations. However, it is usually enough for the evaluator to know that something attendant upon the installation of Harvard Project Physics (and not an extraneous, "uncontrolled" influence unrelated to the curriculum) is responsible for the valued outcomes; to give a more definite answer about what that something is would carry evaluation into analytical research.

5. Autonomy of the Inquiry
The independence and autonomy of science is so important that Kaplan (1964, pp. 3-6) wrote first of it in his classic The Conduct of Inquiry:

It is one of the themes of this book that the various sciences, taken together, are not colonies subject to the governance of logic, methodology, philosophy of science, or any other discipline whatever, but are, and of right ought to be, free and independent. Following John Dewey, I shall refer to this declaration of scientific independence as the principle of autonomy of inquiry. It is the principle that the pursuit of truth is accountable to nothing and to no one not a part of that pursuit itself.
Not surprisingly, autonomy of inquiry proves to be an important characteristic for distinguishing elucidatory and evaluative inquiry. As was seen incidentally in the above quote from Cronbach and Suppes, evaluation is undertaken at the behest of a client who expects that particular questions will be answered. Elucidatory inquiry must be free to follow leads which pique the curiosity of those who know most intimately the aspirations of the discipline.

6. Properties of the Pheonomena Which are Assessed
Evaluation is an attempt to assess the worth of a thing, and elucidatory inquiry is an attempt to assess scientific truth. Except that truth is highly valued and worthwhile, this distinction serves fairly well to discriminate elucidatory and evaluative inquiry. The distinction can be given added meaning if "worth" is taken as synonymous with "social utility" (which is presumed to increase with improved health, happiness, life expectancy, increases in certain kinds of knowledge, etc., and decreases with increases in privation, sickness, ignorance, and the like) and if "scientific truth" is identified with two of its possible forms: 1) empirical verifiability of statements about general phenomena with accepted methods of inquiry; 2) logical consistency of such statements. Elucidatory inquires may yield evidence of social utility, but only indirectly -- because empirical verifiability of general phenomena and logical consistency may eventually be socially useful.

In this view, all inquiry is seen as directed toward the assessment of three properties of statements about phenomena: 1) their empirical verifiability by accepted methods; 2) their logical consistency with other accepted or known facts; 3) their social utility. Most disciplined inquiry aims to assess each property in varying degrees. The definition of "theory" in Webster's Third New International Dictionary, (definition 3.a (2)) is tripartite: "The coherent set of hypothetical, conceptual and pragmatic principles forming the general form of reference for a field of inquiry (as for deducing principles, formulating hypotheses for testing, undertaking actions)." The three inquiry activities in Webster's definition correspond closely to the three inquiry properties proposed here.

7. "Universality" of the Phenomena Studied
Perhaps the highest correlate of the elucidatory-evaluative distinction is the "universality" of the phenomena studied. The term universality is a bit grand, but it conveys a meaning that might have been distorted by many more modest labels that suggest themselves. Elucidatory inquirers work with constructs having a currency and scope of application which make the objects one evaluates seem parochial by comparison. A psychologist experiments with "reinforcement" or "need achievement" which are regarded as neither specific to geography nor to one point in time. The effect of positive reinforcement following upon the responses that are observed is assumed to be a phenomenon shared by most persons in most times; moreover, the number of specific instances of human behavior which are examples of the working of positive reinforcement is great. Not so with the phenomena which educationists evaluate. A particular textbook, an organizational plan, and a film strip have a short life expectancy and may not be widely shared. However, whenever their cost or potential pay-off rises above a negligible level, they are of interest to the evaluator.

Three aspects of the "universality" of a phenomenon can be identified: 1) generality across time (Will the phenomenon -- a textbook, "Self-concept,"etc. -- be of interest fifty years hence?); 2) generality across geography (Is the phenomenon of any interest to people in the next town, the next state, across the ocean?); 3) applicability to a number of specific instances of the general phenomenon (Are there many specific examples of the phenomenon being studied or is this the "one and only"?).

8. Salience of the Value Question
At least in theory, a value can be placed on the outcome of any inquiry, and all inquiry is directed toward the discovery of something worthwhile and useful. In evaluation, it is usually quite clear that some question of value is being addressed. Indeed, value questions are the sine qua non of evaluative inquiry and usually determine what information is sought. This is not to say that value questions are not germane in elucidatory inquiry; they are just less obvious. The acquisition of knowledge of auto-mechanics or the inculcation of "good citizenship" are clearly value-laden endeavors. The value questions in the derivation of anew oblique transformation technique in factor analysis or the investigation of the transfer of information from short-term to long-term memory are not so obvious, but they are there nonetheless. With respect to assessing the value of things, the difference between elucidatory and evaluative inquiry is one of degree, not of kind.

9. Investigative Techniques
A substantial amount of opinion has been expressed recently to the effect that elucidatory and evaluative inquiry should employ different techniques for gathering and processing data, that the methods appropriate to elucidatory inquiry -- such as comparative experimental design -- are not appropriate to evaluation, or that with respect to techniques of empirical inquiry, evaluation is a thing apart. In fact, however, there are many more similarities than differences between elucidatory and evaluative inquiry with regard to the techniques by which empirical evidence is collated and judged to be sound. As Stake and Denny (1969, p. 374) indicated: "The distinction between research and evaluation can be overstated as well as understated. Researchers and evaluators work within the same inquiry paradigm .... [training programs for] both must include skill development in general educational research methodology."

Hemphill (1969, p. 220) expressed the same opinion when he wrote: "The consequence of the differences between the proper function of evaluation studies and research studies is not to be found in differences in the subject of interest or in the methods of inquiry of the researcher and of the evaluator."

RELATIONSHIP BETWEEN ELUCIDATORY AND EVALUATIVE INQUIRY
Evaluation borrows inquiry techniques and knowledge for recognizing value from basic sciences and contributes little in return. Methodological research in the social sciences produces the technologies of data collection and analysis that are so important to empirical educational evaluation. In addition, knowledge produced by basic elucidatory inquiry is often critical in determining whether a particular finding from an evaluative study counts as good or bad. For example, the evaluative meaning given to the finding that a health curriculum decreased the incidence of cigarette smoking among teenagers is dependent upon the extensive medical research which established a casual link between cigarette smoking and cancer and coronary disease.

The view of the relationship between elucidatory and evaluative inquiry presented here is parasitic; some see the two living symbiotically:

To some extent evaluative research may offer a bridge between 'pure' and 'applied' research. Evaluation may be viewed as a field test of the validity of cause-effect hypotheses in basic science whether these be in the field of biology (i.e., medicine) or sociology (i.e., social work). Action programs in any professional field should be based upon the best available scientific knowledge and theory of that field. As such, evaluations of the success or failure of these programs are intimately tied to the proof or disproof of such knowledge. Since such a knowledge base is the foundation of any action program, the evaluation research worker who approaches the task in the spirit of testing some theoretical proposition rather than a set of administrative practices will in the long run make the most significant contribution to the program development. (Suchman, 1967, p. 1970)
Suchman saw "action programs" as based on scientific knowledge and theory; I see a more tenuous link between (a) what can be conceived of abstractly and established empirically in the laboratory and (b) what can be implemented in the field. There can hardly be said to exist an "intimate" connection between the success of a program in the field and some theory or hypotheses from a basic discipline. The tribal shaman may be effective for all of the wrong reasons, and a prototype turbine automobile may fail even though it is consistent in all particulars with physical theory.

One need not be able to explain phenomena to evaluate them. We can know how well or how poorly without knowing why. Understanding or elucidation is required neither in summative nor formative evaluation. We can perfectly well make summative judgments -- and we frequently do -- without understanding why one of the alternatives being evaluated scored highest on some weighted scale of value. In formative evaluation, observations of the quality of performance are made for the purpose of improving the thing that we are developing. It may seem that one can't very well make an improvement in the system unless one knows why performance was unacceptable. Otherwise, responses to data on substandard performance would be random stabs in the dark with little chance of success. If test scores indicate the prevalence of ignorance among tenth graders after studying a program unit on the relationship between birth rates, death rates, and population growth, the curriculum developers will act quite differently if they suspect the failure is due to the lack of practice in working problems rather than the complexity of the mathematics used to explain the relationship. Surely then, formative evaluation requires some greater knowledge of why performance is acceptable or unacceptable than does summative evaluation. But this greater knowledge is knowledge of particulars not codified general knowledge in the form of an abstract system of laws. We know why many things are as they are, even why somethings are good and others bad; and all of this knowledge is of particulars; it is of no general significance. My car continually loses front-end alignment because of the accident that sprung the frame; my electric razor works poorly because I dropped it on the bathroom floor. Both statements are explanations, but that doesn't make them of general interest. We pursue both formative and summative evaluation without the need to explain outcomes by reference to a system of general laws or relationships among variables.

PRETENTIONS TO SCIENTIFIC EDUCATION
The conclusion forces itself upon one that elucidatory inquiry on education over that past eighty years (from Joseph M. Rice to the greenest Ph.D. just now taking final orals) has been a failure for the most part. The areas that I would regard as distinctly educational research and not highly dependent on research in one of the behavioral or social sciences are such as the following: teacher behavior, teacher competence, classroom interaction, guidance theory, counseling, tests and measurement, media of instruction (e.g., television and movies), the organization of teaching (e.g., team teaching, cross-age teaching, and differentiated staffing), class size research, and many of the prominent areas of Ralph Tyler's (1967-68, p. 44) map of science education (the outcomes of science education, student development -- perhaps as opposed to human development in general -- the development of teachers, the objective of education for teachers, the teaching-learning process of teacher education, and the process of change in education, etc.).

I am not under the mistaken notion that scientific inquiry must result in errorless functional relationships among variables. Probabilistic laws are the order of the day among the sciences (physical, natural, or social). Our codified bodies of knowledge in any area are actually tendency statements about the occurrence of phenomenon A tending to be associated with the occurrence of phenomenon B. But educational research has hardly produced even tendency statements about tendency statements, i.e., that one has observed a tendency for A to tend to be associated with B sometimes and under some conditions.

Allan Mazur (1968) gave a wry assessment of the status of sociology vis a vis the more established physical and natural sciences. Mazur used as a touchstone for determining when a body of inquiry had become a science whether the discipline strikes the vast majority of people-in-general as profound. A science, he maintained, "is a body of theoretical knowledge that is not trivial. The scientist must have better theories than the layman or he's really not a scientist at all. If physics is a science because it is empirical and theoretical, it is also a science because physicists' theories about the physical world work better than the non-physicists' theories." "An empirical theoretically connected body of knowledge is science only when the people who know about the theories know more about the real world than the people who do not know the theories." At that point, Mazur asked whether sociology is a science, and answered, No. Perhaps the same frame of mind led a well-known philosopher of science to remark at lunch one day that John McDonald, author of the Travis McGee detective story series, is a better sociologist than most contemporary sociologists. Walter Lippman, who seldom revealed his academic credentials (which as a student of William James and George Santyana were considerable), made the same point several years before Mazur about the psychology of the1920's:

"We can be confident that on the whole a good meteorologist can tell us more about the weather than even the most weather-wise old sea captain. But we cannot have that kind of confidence in even the best of psychologists. Indeed, an acquaintance with psychologists will, I think, compel anyone to admit that, if they are good psychologists, they are almost certain to possess a gift of insight which is unaccounted for by their technical apparatus. Doubtless it is true that in all the sciences the difference between a good scientist and a poor one comes down at last, after all the technical and theoretical procedure has been learned, to some sort of residual flair for the realities of the subject. But in the study of human nature that residual flair, which seems to be composed of intuition, commonsense, and unconsciously deposited experience, plays a much greater role than it does in the more advanced sciences." (Lippman, 1929, p. 162)
The "science of education" certainly fails in its pretensions to scientific status if we use Mazur's criterion. There is probably far more knowledge (about how to manage and promote learning in the classroom) in the nervous systems of ten excellent teachers than an average teacher can distill from all the educational research journals in existence. Hawkins (1966) wrote that in teaching and learning "... the best practice excels the best theory in quite essential ways ...."

Teacher behavior studies provide a good illustration of the unproductive nature of most elucidatory educational inquiry. The analogy is uncomplimentary, however true it is, that most teacher behavior studies, indeed, even the best of them, have found relationships no stronger than the evidence for the ability of graph-analysts to assess personality through handwriting samples. The data in Table 1 show a typical graph-analyst's ability to rate personality traits from hand-writing samples to be nearly as good as many judge's ability to rate teacher behavior characteristics and the subsequent relation of these characteristics to the gains made by students on achievement tests. In Table 2, correlations are shown between teacher behaviors (based on the works of Bellack, Smith, Taba, Flanders, and others) and pupil achievement from one of the best studies ever performed in this area. The correlations are not indicative that instructional researchers' understanding of these phenomena is conspicuously superior to graph-analysts insight into personality.


The study of teacher behavior and its relation to student performance may be even worse off than my impertinent analogy would indicate. Barak Rosenshine has recently reviewed the experimental literature growing out of the correlational teacher behavior studies and concluded that the results of these experimental studies are greatly discouraging. When a variable once shown correlationally to have a weak or moderate relationship to student performance as manipulated experimentally, in almost all cases no influence on student performance was observed. Such is the residue of some hundreds of thousands of elucidatory inquiry about the dominant personality in the classroom.

Certainly one of the most important factors militating against successful elucidatory inquiry about education is the enormous complexity of the system educational researchers seek to understand. This reason is often advanced by educational researchers themselves, and in advancing it they bring down upon themselves accusations that they are merely making excuses for lack of diligence and insight. I certainly do not make such an ungenerous assessment of this explanation for the failure of a great deal of educational research.The system educational researchers study is as complex as any that science has ever dealt with successfully; a lake or the human heart is trivial by comparison. But the fact remains that elucidatory inquiry on education simply does not seem to have turned up any important, reliable, replicable relationships worthy of continued study. Several of these will be needed before any body of theory or even a modestly complicated model of education can be constructed. Until we get these fundamental relationships our pretensions to elucidatory inquiry on education will be vain boasting.

THE PROBLEM OF NON-GENERALIZABLE KNOWLEDGE
Psychology has the phenomenon of reinforcement which works fairly reliably. The fascinating thing about Piaget's work on cognition is that it demonstrates such a stable, replicable phenomenon: Take any three-year-old off the streets of Cleveland or out of the rain forests of Brazil, and that child won't conserve mass. But the phenomena and facts that educational research has unearthed are so fragile as to scarcely deserve the name "fact." Interactions predominate in the accumulated wisdom of educational research. A relationship among variables appears and disappears so often as a function of extraneous conditions that one genuinely never knows when to expect to see the relationship.

It has been observed that a young person seeking to make their reputation as a scholar could hardly do better than choose educational research. Current findings in the field are so interactive and ungeneralizable that a novice need only stake out a small area of interest and apply themself diligently to studying it. They will surely rise to world-wide eminence confident that their colleagues in related areas will never devise models or theories sufficiently general either to subsume their work or to prove it wrong.

The most useful scientific laws are those of physics, most of which are learned firsthand and nearly unconsciously early in life by everyone. These basic physical laws are nearly perfectly generalizable; there is no evidence that they have ever been suspended on the surface of the earth at any time in history. An astronaut in a space-capsule experiences the repeal of certain physical laws and must laboriously learn a new set. By contrast, the laws of the social and behavioral sciences are of extremely limited generality. These laws are highly interactive with factors such as geography, time, culture, and characteristics of persons. If physical Laws were as limited in generality as the laws so far discovered by social scientists, we would hesitantly creep out of bed each morning not knowing whether we would float to the ceiling or crash to the floor. If physical Laws were as erratic as the "laws" governing the educational system, we wouldn't dare to get out of bed.

Consider one of the most interesting areas of elucidatory inquiry in education in recent years: the work of Rothkopf, Frase and Anderson and others on mathemagenic behavior, in particular the control of attention. These researchers have demonstrated the striking effects on learning of the control of attention through prompts in programed instruction and adjunct questions in textual materials. On the latter question, it has been demonstrated that questions preceding paragraphs to be read cue the learner's attention to the answers to the questions and away from incidental material. Post-paragraph questions cannot cue specific passages in a paragraph and thus result in greater acquisition of information. In programed instruction, over-prompting of stimulus and response in associative learning has been shown to cause inattention to the stimulus-response pairing resulting in poor learning.

The conclusions these researchers have reached concerning the cueing and control of attentional processes are as well understood and lawful as basic educational research can hope for. Yet further work in this area has shown the facilitative effect to be highly interactive with such factors as positioning and pacing of questions, the difficulty of the material to be learned, the propinquity of question and answer, whether questions are general or specific, whether answers to questions are common or uncommon words, whether subjects can turn back in a text for review, habituation to the adjunct questions, the motivation of the learner, amount of rewards for learning, meaningfulness of material to be learned, immediate or delayed retention, the mode (visual or auditory) or presentation of the material, etc. The mathemagenic effect appears and disappears as a function of these mediating conditions even in the scientists' own make-believe world of ceteris paribus, far away from the welter of variables in a typical classroom.

J. M, Stephens in The Process of Schooling (1967) surveyed a huge portion of the educational research literature. He sampled the more trustworthy regions of that literature and came up with a distressing showing of irregularity in the findings of educational researchers for the past sixty years. Stephens's argument is that in a dozen or more areas of the teaching-learning domain, the studies investigating a certain relationship between variables divide about equally between positive and negative findings. I interpret Stephens's summary to show again that relationships among variables in education are highly interactive. To put the matter in slightly different language, the relationships which educational researchers investigate are highly context dependent. A relationship between televised instruction and student learning appears here because of administrative support and fails to appear there because of teacher attitude. In other words, the relationships among educational phenomena are very likely mediated by a huge number of influences, most of which we have yet to conceptualize properly, let alone measure validly and reliably.

If educational researchers would strive for generalizability of findings, by attempting to relieve context dependence, I am convinced that their inquiries would be pushed further and further toward the psychologists' laboratory. If elucidatory inquiry in education could be made less and less context dependent, it would look more and more like basic psychological inquiry.

IS THE GAME WORTH THE CANDLE?
We are heading here straight toward the question of whether education should persevere in attempts to build a science of education, that is, in attempts to build codified bodies of knowledge leading to models and ultimately theories about the teaching-learning and broader educational process. In a recent interview concerned with the prospects of educational research, the philosopher of science, Thomas Kuhn, seemed to answer the question just posed in the negative:

I'm not sure that there can now be such a thing as really productive educational research. It is not clear that one yet has the conceptual research categories, research tools, and properly selected problems that will lead to increased understanding of the educational process. There is a general assumption that if you've got a big problem, the way to solve it is by the application of science. All you have to do is call on the right people and put enough money in and in a matter of a few years, you will have it .But it doesn't work that way, and it never will." (See Dershimer (1970, p. 79).
I interpret Kuhn's remark to imply that the construction of a science of education is not impossible, but it is scarcely the simple matter of a redoubling of effort and faith in the application of the scientific method. The opinion is frequently expressed that educational research can never be basic research because education is a practice unlike psychology and biology which are "basic." Such reasoning is mere word play, not unlike arguing that basic research in physics is impossible because a carpenter practices physics when setting a screw (inclined plane) or pulls a nail (lever and fulcrum). Ebel (1967, p. 81) gave as one of his reasons that "basic research in education can promise very little improvement in the process of education" that "the process of education is not a natural phenomenon of the kind that has sometimes rewarded scientific study in astronomy, physics, chemistry, geology, and biology. "Ebel's selection of examples deftly excluded all of the social sciences. One wonders whether he would regard them as studies of natural phenomena. The English jurist Glanville Williams dispensed with arguments from nature against contraception by saying that "the statement that an act is unnatural coming from a moralist means little more than that he doesn't like it." Ebel's use of the word appears to mean little more than that he would prefer to exclude education from basic inquiry. I do not. I believe that we could pursue with success the construction of a science of education. This pursuit could culminate in non-trivial theories with explantory and predictive power. But the question before us now is should we? Should we do it now and is it worth the price?

In my opinion, we should not. Attempts to build a theory of science teaching, for example, face the problem of a rapidly changing object of study. The phenomena that physicists have attempted to explain have a certain timeless quality about them: force, velocity, mass, etc. To be sure physicists probe new phenomena decade after decade. However, they have confidence that the stuff of which physical studies are made, will be around for a long time. In the social sciences one finds far less stability in the subject matter. Inquiries can become obsolete; circumstance of society change. John Kenneth Galbraith pointed out quite clearly in The AffluentSociety how industrialization and more sophisticated technology upset classical relationships among supply and demand and changed the function of the market. The change in the phenomena of interest to educational researchers is even more chaotic because educators pick up and discard one new idea after another. To the extent that researchers have extracted something stable and of interest out of this welter of educational practices, they look and act more like learning psychologists than like scientists studying the education system. Published educational research has an aroma of Chippendale about it. It is generally dated as compared with the concerns of contemporary schools (as witness research on handwriting, "study hall," instructional radio, progressivism, teaching machines, attendance, et al., which once seemed so timely and today seem as quaint as the Packard, Allen's Alley, or Bill Haley and the Comets). While the researcher is groping for understanding of the effects of televised instruction in large groups, the schools will long since have moved on to dual audio television or video-type cassettes embedded in individualized instruction.

The problems of obsolescence of educational research findings would be aggrevated if researchers followed some of the more cogent suggestions for building a science of teaching-learning. Jack Easley, a former colleague of mine at the University of Illinois, maintained that there was some hope for a modest science of education as long as the models or theory included learner, teacher, and subject matter as elements. He felt that attempts to theorize in general without regard for the structure of the subject matter to be taught would flounder when applied to particular instances of teaching. Thus we might successfully pursue the construction of a theory of teaching the concept of natural selection to post-adolescents, but anything as general as a theory of teaching science to human beings was certain to fail. I regard Easley's position as wise counsel. Its implications for the possibility of building a science of education are dire, however. For what is more fluid and changing than the concepts and subject matter that each new generation of educators feels must be taught in the classroom. Recently a biologist in charge of one of those massive undergraduate general biology courses told me that in reviewing a 1940s vintage college biology text he discovered virtually no content he would want to teach today.

It can be argued -- and not easily refuted -- that educational change is so chaotic because innovations are not based on scientific knowledge. This position implies that the pace of change would diminish if education rested on a foundation of reliable basic knowledge of teaching-learning. A little knowledge from basic elucidatory inquiry might stabilize the educational system and thus permit more extensive elucidatory inquiry. This position probably attributes too much power to scientific knowledge as a controlling force of a social institution.

My pessimism for elucidatory inquiry in education stems not so much from a lack of faith in our intelligence or ability to pursue scientific inquiry about anything we choose, as it does from an assessment of the resources one could reasonably assume that society could spare for the task of building a science of education. In his proposal for the National Institute of Education, Roger Levien reported on the manpower available fo rresearch, development and innovation in education. In fiscal year 1968, the total effort from all sources (public, private, governmental, etc.) for research on education was 1,930 person-years. The definition of research in that report included both basic and applied, and hence was construed more broadly than I have used the term "elucidatory inquiry" here. We can safely conclude, then, that as recently as three years ago the total effort directed toward the construction of a systematic body of knowledge about education amounted to substantially less than four thousand persons in the United States working half days. By contrast, in the same year, 1968, nearly 15,000 full-time researchers were grappling with the problems of agricultural research, some fraction of them no doubt building a science of agronomy. In the health sciences nearly 60,000 person-years were expended on research and development in fiscal year 1968. The poverty of the educational research effort is more than we might have expected. The construction of a science of education might drag on for many decades if it depended solely on the efforts of 1,930 full-time "educational scientists." In short, the resources for inquiry of any type, elucidatory or evaluative, in education are severely limited, will probably remain so for some time, and perhaps ought not be squandered on such a dubious endeavor as building a science of education. Scriven (1969, p. 19) remarked in the context of a discussion of pure and applied research in psychology that "it is much better to move into theory exactly where and only when obliged to by the combination of data and needs that define our tasks.Speculation in the absence of a clarification of these parameters is too often merely idle, the kind of irresponsible gambling with society's resources that is lauded in cheap histories of science and was once a rich amateur's prerogative."

We can afford not to seek a theory of education at this time simply because we have always lived without complete understanding. Indeed it is usually better and safer at the level of the whole society to know how well than to know why, if we must choose one or the other. People do many things better than they ought to, e.g., football, trumpet playing. We can do many things better than generalizable codified knowledge about those things would seem to permit. A great deal of our knowledge is inexpressible, yet it exists and is important nonetheless. Anthropologists have discovered primitive societies that are ignorant of the relationship between sexual intercourse and pregnancy, yet presumably the existence of these societies is no more precarious for that ignorance.

Persons with a taste for empirical, disciplined inquiry on science teaching have more serious business to undertake than elucidatory inquiry on education. Their attentions ought to be focused on educational development and evaluative inquiry. Such developments in science education and elsewhere might best be regarded as bearing a heuristic relationship to elucidatory inquiry in the social sciences, particularly the science of human learning. I have developed this notion of a heuristic relationship between basic and applied knowledge to some length elsewhere and space does not permit repeating it here (see Glass, 1971). Perhaps it is sufficient to say only that the word "heuristic" is used in the literal sense, namely suggesting or stimulating empirical research. Basic knowledge will never prescribe a particular practice, but it can stimulate creative minds who have some understanding of that basic knowledge to develop materials and practices based loosely on it. Basic knowledge out of psychology and other social and natural sciences ought to be mediated through the minds of the one in ten thousand masterful teachers (the Gattegnoes, Papiis, and Ashton-Warners) currently at large who would create curriculum materials under the inspiration of this knowledge. My model of educational development involves inspiring these creative geniuses with what reliable knowledge exists in the basic sciences that is conceivably relevant to education, allowing these teachers to create developments, and then subjecting their creations to the reality testing of evaluation. Unfortunately there is not space here to develop these ideas fully. Some of the flavor of my position can be sensed by comparing and contrasting it with positions that have been published. My assessment of the pay-off of elucidatory inquiry in education agrees totally with Atkin's (1967-68) assessment of "behavioral science" research on science teaching. I am more taken with the "engineering model" of development than he and I resonate to the tones of his eloquent description of the "naturalistic model" of research and development, though the resonance is dampened somewhat by the same intolerance of ambiguity that makes the "engineering model" attractive to me.

There is much work to be done in the development of evaluation methodology in education. An enormous amount of work remains to be done in the measurement of human behavior. I am encouraged by the amount of effort in that direction that researchers on science teaching have made. The results, as they appear in the Journal of Research in Science Teaching are increasingly impressive. For the next several decades then, I am placing my money on evaluative inquiry (with a heavy emphasis on measurement) which is integrated with the process of educational development in which master teachers play an important role. As for elucidatory inquiry about science teaching or education generally, I am far less optimistic.

REFERENCES

Atkin, J. Myron. (1967-68). Research Styles in Science Education. Journal of Research in Science Teaching, 5, 338-345.

Braithwaite, Richard B. (1953) Scientific Explanation. Cambridge, England: Cambridge University Press.

Cronbach, Lee J. and Suppes, Patrick. (1969). Research for Tomorrow's Schools: Disciplined Inquiry for Education. New York: MacMillan.

Dershimer, Richard A. (1970). The Educational Research Community: Its Communication and Social Structure. Bureau of Research, U.S. Office of Education, Project No. 8-0751.

Ebel, Robert L. (1967, October). Some Limitations of Basic Research in Education. Phi Delta KAPPAN, 81-84.

Glass, Gene V. (1971). Educational Knowledge Use. Educational Forum (in press).

Hawkins, David. (1966). Learning the Unteachable. In Shulman, L. and Keislar, E. (Eds.) Learning by Discovery: A Critical Appraisal. Chicago: Rand McNally.

Hemphill, John K. (1969). The Relationship between Research and Evaluation Studies. In Ralph W. Tyler (Ed.), Educational Evaluation: New Roles, New Means. The 68th Yearbook of the National Society for the Study of Education, Part II. Chicago: National Society for the Study of Education.

Kaplan, Abraham. (1964). The Conduct of Inquiry. San Francisco: Chandler.

Lippman, Walter. (1929). A Preface to Morals. New York: MacMillan.

Mazur, Allan. (1968). The Littlest Science. The American Sociologist, 3, 195-200.

Pella, M. O. (1966). A Structure for Science Education. Journal of Research in Science Teaching, 4, 250-252.

Suchman, Edward A. (1967). Evaluative Research. New York: Russell Sage Foundation.

Scriven, Michael. (1958). Definitions, Explanations, and Theories. In Herbert Feigel, Michael Scriven, and G. Maxwell (Eds.), Minnesota Studies in the Philosophy of Science, Vol II. Minneapolis, Minnesota: University of Minnesota Press.

Scriven, Michael. (1969). Psychology Without a Paradigm. In Louis Berger (Ed.), Clinical-cognitive Psychology; Models and Integrations. Prentice-Hall.

Scriven, Michael. (1967). The Methodology of Evaluation. In Robert E. Stake (Ed.), AERA Monograph Series on Curriculum Evaluation, No. 1. Chicago: Rand McNally.

Stake, Robert E. and Denny, Terry. (1969). Needed Concepts and Techniques for Utilizing More Fully the Potential of Evaluation. In Ralph W.Tyler (Ed.), Educational Evaluation: New Roles, New Means. The 68th Yearbook of the National Society for the Study of Education, Part II. Chicago: National Society for the Study of Education.

Stephens, J. M. (1967). The Process of Schooling: A psychological examnination. New York: Holt, Rinehart & Winston.

Tukey, John W. (1960). Conclusions vs. Decisions. Technometrics, 2, 423-433.

Tyler, Ralph W. (1967-68). Resources, Models and Theory in the Improvement of Research in Science Education. Journal of Research in Science Teaching, 5 43-51.

Sunday, June 30, 2024

"Pull Out" in Compensatory Education

"PULL OUT" IN COMPENSATORY EDUCATION
Gene V Glass
Mary Lee Smith
Laboratory of Educational Research
University of Colorado

Paper Prepared for the Office of the Commissioner
U.S. Office of Education, November 1977

ACKNOWLEDGMENTS
We were asked to address the problem that forms the subject of this paper by Marshall Smith and Fritz Edelstein of the Office of the Commissioner, U.S. Office of Education. We have no idea whether or not they had opinions about the "pull out" issue at the outset. If they did, they hid them from us. They encouraged us in every way to examine the question freely and render an independent opinion. That requires commendable nerve, even when the stakes are only $3,000.

We gathered a large number of documents and interviewed many people in just a few weeks. It would not have bean possible to produce this report without the voluntary cooperation of dozens of persons who consented on a moment's notice to be interviewed, or dig up old data, or mail us their reports and papers. The persons who gave us interviews are acknowledged in the Appendix. The staff of System Development Corporation, particularly Ralph Hoepfner and Clarence Bradford, were helpful in performing new data analyses and explaining old ones. Len Cahen of the Far West Laboratory for Educational Research and Development shared several unpublished papers with us. Bob Stonehill of the U.S. Office of Education found reports for us that we wouldn't have obtained otherwise.

SUMMARY
This report examines the research on "pull out," a method or type of school organization for remedial teaching of Title I eligible pupils. Four major issues addressed are: (1) the educational benefits of pulling students out of the daily routine to provide them with compensatory education services; (2) the impact of such action on students; (3) whether a child is better served if he remains in the classroom all day; and (4) alternatives to "pull out" available for providing compensatory assistance to educationally disadvantaged children. Other related issues examined are: the prevalence of "pull out" programs, the benefits or losses resulting from "pull out" programs, teacher contact with and attitudes toward pulled out pupils, financial costs of "pull out" programs, and the potential contribution of "pull out" programs to cultural separatism, racial segregation, or even racism. It is concluded that despite the near universality of pulling Title I eligible pupils out of regular classrooms for compensatory instruction, the procedure has neither academic nor social benefits, may be detrimental, and is used mostly to satisfy Title I regulations. Alternatives to "pull out" are recommended.

"Pull Out" is a method or type of school organization for remedial teaching of Title I eligible pupils. With this plan, Title I eligible pupils are pulled out of regular classes containing both eligible and non-eligible pupils and sent to a different room to receive instruction from a remedial specialist teacher. "Pull Out" has emerged as a prominent feature of compensatory education in the past few years, and it now concerns policy-makers, researchers, and educators alike. This report was written in response to a request from the Office of the Commissioner of Education to examine the research on "pull out." In the course of preparing this opinion, we interviewed about thirty persons in schools, state education agencies, the federal government, universities, and teacher organizations; in addition, we read and, in some instances, reanalyzed data from approximately 150 documents.

The Incidence and Context of "Pull Out"
Roughly 75% of compensatory education pupils receive remedial reading instruction in the "pull out" setting; the comparable figures for mathematics and language are 45% and 41%, respectively. When these figures are corrected to eliminate pupils in 100% Title I eligible classrooms who do not need to be "pulled out," the "pull out" rates in all other classrooms rise to 84% for reading, 54% for mathematics, and 50% for language arts. When one considers further that pupils might be "pulled out" for one of these subjects and not the other, it is plausible to say that in classes not 100% "Title I eligible" the practice of "pull out" for compensatory teaching is nearly universal.

"Pull out" is probably more prevalent in small and medium-sized districts than in large, urban schools. "Pulled out" and "mainstreamed" pupils compensatory education pupils taught in regular classes do not differ in their academic performance before remedial teaching; thus, "pull out" seems not to be prescribed differentially for pupils with varying remedial needs.

The amount of the entire instructional day spent in the "pull out" setting rose from around 5% in 1973-1974 to around 9% in 1974-1975. The percentage of time does not vary by subject taught (reading vs. mathematics) or by grade level (elementary vs. secondary). Although time in the "pull out" setting is a small part of the total instructional day, at Grade 1 it constitutes almost half of the instructional time funded by TitleI. At the elementary school level, one-fifth of the "pulled out" pupils miss regular classroom instruction in the subject for which they are removed from the regular class (i.e., they are "pulled out" of regular reading to receive remedial reading). One-fourth miss social studies; one-seventh miss science. One-third miss no academic subject at all since they are pulled out during study periods in the regular class. (By some chop logic we do not understand, supplanting is not supplanting at all if one supplants science and social studies.)

The "pulled out" pupil has three chances in four of receiving remedial instruction from a remedial subject matter specialist. (The comparable chances for a mainstreamed compensatory education pupil are only one in three.) However, the "specialist" teachers receive very little training for their job (less than ten hours in any one year on the average) and they receive virtually no extra pay (less than 5% more than regular teachers). These data seem to indicate that remedial specialists are distinguished from regular teachers neither by more intensive training nor by the pay they receive. The most cynical assessment of their role and contribution would be that remedial, specialist teachers are merely rechristened regular classroom teachers -- the motive for so designating them being, perhaps, the need to comply with certain Title I regulations.

Finally,"pull out" programs appear to be roughly twice as expensive per pupil as mainstream compensatory programs, probably because of much smaller class size in the former than the latter.

The Effects of "Pull Out"

Experimental evidence is skimpy on the effects of the "pull out" technique per se on pupils' academic progress. A study recently published by the National Institute of Education (September 1977) alleges to show beneficial effects of "pull out" at certain grade levels and in certain subjects and detrimental effects elsewhere. We have examined the data and find little support for these conclusions. The academic gains made by "pulled out and "mainstreamed" compensatory education pupils in the NIE data are virtually identical, differing overall grades and subjects by less than one-quarter month in grade equivalent units. Perhaps a better database for assessing the effects of "pull out" exists in the data files of the evaluation of the Emergency School Aid Act (ESAA) conducted by Systems Development Corporation. There one finds a consistent negative relationship between the percentage of time pupils spend in the "pull out" setting and their math and reading achievement. This relationship was consistent across all grades and subjects; it held true for samples numbering about 10,000 pupils in total. Moreover, the relationship persisted even after more than a dozen background variables were controlled statistically.

A vast body of empirical research on instructional methods and organization is pertinent to the "pull out" problem because the phenomena investigated share various features with the "pull out" technique. Such related topics include the following: (a) ability grouping, (b) mainstreaming the handicapped, (c) racial desegregation, (d) labeling pupils with consequent changes in teachers' expectations of them, and (e) peer tutoring.

Our synopses of the research evidence on these topics are as follows: (a) The research on ability grouping is inconsistent, uninformative, and a battleground for various social ideologies; it was not helpful to us in forming opinions about "pull out." (b) The research on mainstreaming the handicapped was exceedingly skimpy, but the findings of three studies point toward the benefits of integrating EMH, EMR, or emotionally disturbed pupils into the regular classroom. (c) Nearly all research on racial desegregation fails to trace racial mixing at levels lower than the school building. Using the Coleman data, McPartland (1969) assessed the effects on verbal achievement of black pupils having predominantly black classmates instead of white classmates. Although the effect diminished as more background variables were partialed out, the effect of racial segregation at the classroom level was negative in first analyses and never appeared beneficial regardless of how many variables were statistically controlled. (d) Research into the effects of labeling pupils on teachers' behavior toward them and expectations of them proved to be most pertinent and startling. Labeling a pupil "mentally retarded," "intellectually slow," or "academically weak," reduces his academic achievement by one-quarter standard deviation below that of comparable pupils not so labeled. Furthermore, teachers' attention and support for pupils invidiously labeled are reduced by one-third standard deviation below those for comparable unlabeled pupils; and teachers' judgment of labeled pupils' success, motivation, social competence, etc. is reduced by nearly one-half standard deviation. These findings from more than forty experiments indicate that the effects of labeling pupils are large and worrisome. (e) Finally, research on peer tutoring, which presumably could occur less often in the "pull out" programs. Pupils pulled out of regular classrooms would have to receive remarkably effective compensatory programs to offset the potential risks incurred. In our opinion, the "pulled out" pupil is placed in moderate jeopardy of being dysfunctionally labeled, of missing opportunities for peer tutoring and role modeling, and of being segregated from pupils of different ethnic groups.

Historical and Political Context of "Pull Out"

We believe that the "pull out" problem was created by the ESEA Title I regulations and the manner in which they have been interpreted and enforced. To quote one state education department official, "'Pull out' exists for one reason only; because the 'locals' are afraid Big Brother will catch them in a 'supplanting' violation." The practice of pulling Title I eligible pupils out of regular classrooms so that a "specialist" teacher could give them instruction in a separate classroom did not grow out of professional judgment about curriculum or instruction. The history is complex, but nearly any disinterested reading of it leads to the same conclusion: "pull out" is an artifice created by schools at the urging of USOE's minions in state education departments to satisfy regulations concerning "supplementing, not supplanting" and "excess costs." The regulations themselves reflect a philosophy that seems seldom to have been seriously challenged. They are enforced with a nearly obsessive concern that "noneligible" pupils might receive Title I services. Yet the argument can be made that even pupils performing at grade level and above are educationally deprived by merit of attending a school with large concentrations of poor (since such schools attract less qualified teachers, have poorer opportunities for peer tutoring, etc.).

One finds virtually no support for the "pull out" concept among educators or their professional organizations. Teachers worry that pulling pupils out of class creates discontinuities in their schooling and makes coordination of teaching difficult. Others worry that the regular classroom teachers will feel less responsible for pupils whose needs are presumably being met somewhere else by a specialist teacher. The National Education Association regards "pull out" as a minor issue and will merely watch its evolution, being concerned only with keeping pupil-to-teacher ratios low. The "pull out" problem seems to be no one's major concern. But it may well be one of those quiet, inconspicuous matters that count heavily in ways seldom clearly seen.

Conclusions, Observations, and Recommendations

Our work has led us to the following conclusions and observations about the "pull out" technique and several recommendations for dealing with the problems it raises.

  1. Pulling Title I eligible pupils out of regular classrooms for compensatory instruction is virtually universal.
  2. The "pull out" procedure per se has no clear academic or social benefits and may, in fact, be detrimental to pupils' progress and adjustment to school.
  3. The "pull out" procedure is used by schools more to satisfy Title I regulations than because it is judged by teachers to be a sensible and beneficial plan.

We wish to bring the following recommendations to the attention of those persons at all levels who administer Title I programs and who will influence the evolution of compensatory education:

  1. The Title I regulations, which now reflect an overweening concern with targeting funds on "eligible" pupils, should be examined.New considerations should be given to the needs of all pupils in poor schools and the integrity of total school programs.
  2. Instructional strategies should be designed that would eliminate the invidious labeling of compensatory education pupils and their segregation from classes of "regular" pupils.
  3. Teachers, administrators and other persons connected with Title I programs should be informed of the findings of research on the "pull out" method and associated phenomena. 4.Methods should be devised of counteracting the possibly detrimental effects of "pull out" where educators choose to use it or have no reasonable alternatives. Such methods could include means for coordinating instruction across two sites and techniques of teacher observation that lessen the possibility that "pulled out" pupils will be unconsciously neglected in regular classes.

BIBLIOGRAPHY

Bailey, S.K. and Mosher, E.K. (1968). ESEA: The Office of Education Administers a Law . Syracuse, N.Y.: Syracuse University Press.

Billett, R.O. (1932). The administration and supervision of homogeneous grouping. Columbus:The Ohio State University Press.

Borg, W.R. (1966). Ability Grouping in the Public Schools, (2nd ed.). Madison, Wis.: Dembar Educational Research Services, Inc.

Campbell, E.Q. (1965). Structural effects and interpersonal relationships. American Journal of Sociology, 71, 284-289.

Cornell, E.L. (1936). Effects of ability grouping determinable from published studies. In G.M. Whipple (Ed.),The Ability Grouping of Pupils. Yearbook of Nattional Social Studies Education, Part I. Bloomington, Ill.: Public School Publishing Co., pp. 289-304.

Coulson, J.E., et al. (1977). The Third Year of Emergency School Aid Act (ESAA) Implementation. Santa Monica, Calif.: Systems Development Corporation.

Dienemann, Paul F., Donald L. Flynn, and Nabeel Al-Salam. (1974). An Evaluation of the Cost Effectiveness of Alternative Compensatory Reading Programs. Bethesda, MD: RMC Research Corporation.

Eash, M. (1961). Grouping: What have we learned?. Educational leadership, 18 429.

Ekstrom, R. (1961). Experimental studies of homogeneous grouping: A criticalreview. School Review, 69, 216-226.

Esposito, D. (1973). Homogeneous and heterogeneous ability grouping: Principal findings and implications for evaluating and designing more effective educational environments. Review of Educational Research, 43 (2), 163-179.

Findley, W.G. and M.M12\Byran.Ability Grouping: 1970,Status Impact and Alter- natives.Athens, Georgia:University.of Georgia, Centerfor.EducationImprovement, 1971. \

Flynn, D.L, Hass, A.E., & Al-Salam, N.A. (1976). An Evaluation of the Cost Effectiveness of Alternative Compensatory Reading Programs. Bethesda, Maryland: RMC Research Corporation.

Gampel, D.H., Gottlieb, J., and Harrison, R.H. (1974). Comparison of classroom behavior of special-class EMR, integrated EMR, low IQ, and nonretarded children. American Journal of Mental Deficiency, 79, 16-21.

Glass, Gene V, et al. (1970). Data Analysis of the 1968-1969 Survey of Compensatory Education. University of Colorado.

Glass, G.V (1977). Integrating findings: The meta-analysis of research. Review of Research in Education, (in press).

Goldberg, M.L., Passow, A.H., and Justman, J. (1966). The Effects of Ability Grouping. New York: Teachers College Press, Columbia University, 1966.

Goodman, H., Gottlieb, J., and Harrison, R.H.Social acceptance of EMRs integrated into a nongraded elementary school.American Journal of MentalDeficiency, 1972, 76, 412-413.

Kiesling, H.J. Input and Output in California Compensatory Education Projects. Santa Monica, Calif.:Rand Corporation, 1971.

McDill, E.L., Meyers, E.D., and Rigsby, L.C. Sources of Educational Climates in High Schools. Washington, D.C.: U.S. Office of Education Cooperative,Research Report No. 1999, 1966.

McLaughlin, M. W. Evaluation and Reform: The Elementary and Secondary EducationAct of 1965, Title I. Cambridge, Mass.: Ballinger Publishing Co., 1975.

McPartland, J.M. (1969). The relative influence of school desegregation and of classroom desegregation on the academic achievement of ninth-grade Negro students. Journal of Social Issues, 25, 93-102.

Miller, W.S. and Otto, H.J. (1930). Analysis of experimental studies in homogeneous grouping. Journal of Educational Research, 21, 95-102.

National Institute of Education. The Effects of Services on Student Development. Washington, D.C.: The National Institute of Education, U.S. Department of Health, Education and Welfare, September 1977.

Rock, R.T., Jr.A critical study of current practices in ability grouping.Education Research Bulletin.Catholic University of America, Nos. 5 and 6, 1929.

Smith, M.L. Meta-analysis of teacher expectation research. Unpublished paper. Boulder, Colorado: Laboratory of Educational Research, University ofColorado, 1977.

Smith, M.S. "Equality of educational opportunity: The basic findings reconsidered." In Equality of Educational Opportunity, edited by F. Mosteller and D.P. Moynihan. NY: Random House, Inc., 1972.

Tuckman, B.W. and Bierman, M. Beyond Pygmalion: Ability group reassignment and its effects. Paper presented at the Annual Meeting of the American Educational Research Association, New York, 1971.

Vacc, N.A. A study of emotionally disturbed children in regular and special.classes. Exceptional Children, 1968, 197-204.

Appendix

PERSONS INTERVIEWED ON THE "PULL OUT" ISSUE

Dr. Richard Cortright National Education Association
Dr. George Cronk New York Department of Education
Dr. Joy Frechtling National Institute of Education
Dr. Gerald. Freeborn New York Department of Education
Dr. John Garrett Denver Public Schools
Dr. David Gordon California Department of Education
Dr. Susan S. Hartley Northwest Missouri State College
Office of Senator Floyd Haskell Denver, Colorado
Dr. Ralph Hoepfner Systems Development CorpOration
Ms. Linda Jones Colorado Department of Education
Dr. Martin kaufman BEH, U.S. Office of Education
Dr. Michael Kean Philadelphia Public Schools
Dr. Bernard McKenna National Education Association
Dr. Richard Mallory National Education Association
Dr. Robert Mendro Dallas Public Schools
Dr. Lynn Morris Center for the Study of ':valuation, UCLA
Dr. Iris Rothberg National Institute of Education
Ms. Ann Rutherford Denver Public Schools
Dr. Robert Stonehill OPBE, Office of Education
Dr. Gary Toothaker Superintendent of Rifle Public Schools
Dr. Bruce W. Tuckman Rutgers University
Dr. Jean Wellch Systems DevelopMent Corporation
Dr. David E. Wiley CEMREL

Friday, June 28, 2024

Berliner, D. C. & Glass, G. V (2024) Trust but Verify.

2024

Trust But Verify

 

David C. Berliner and Gene V Glass

Arizona State University

 

School improvement programs that work in some places sometimes don’t work elsewhere. School improvement programs that work with some students may not work with others. Programs that appear to have positive effects in the hands of some teachers may fail to produce good effects with other teachers. If this were not the reality of school improvement, we would have found and implemented excellent programs for every state, district, and classroom in the United States by now. But we haven’t, not by a long shot. Instead, we are continually puzzled as we search for high quality education programs that consistently benefit rural white students, or urban black students, or English language learners from hundreds of nations. We also have problems educating the privileged youth of America’s upper-class communities. The education of children who suffer from “affluenza” (Fernandez & Schwartz, 2013) is as disappointing to many educators as is the slow progress of America’s poor students.

 

 It’s past time to lay aside the belief that what works in one setting with one teacher at one time is very likely to work in another setting with another teacher at another time. Education, says our colleague Lenay Dunn (Berliner, Glass, & Associates, 2014), is a complex, intricate endeavor that entails circumstances we can’t control (e.g., family wealth, parents’ education, community support, and special needs of children), influences we can’t easily identify or measure (such as competing school and district initiatives, classroom culture, peer influence, teacher beliefs, and principal leadership), and results we can neither predict nor easily measure (such as resilience, grit, practical intelligence, social intelligence, and creativity). The complex character of teaching children various subjects limits our ability to design programs that function well wherever they are implemented.

 

 However, one must not despair in the face of this reality. Instead, we should feel privileged that we work in a field that is more complex, and thus more challenging, than physics or rocket science. The late, great economist Kenneth Boulding once remarked that if physical systems were as complex as social systems, we would creep hesitantly out of bed each morning, not knowing whether we were about to crash to the floor or float to the ceiling. Educators face the challenges of these unpredictable social systems every day.

 

 Three Obstacles to Transfer

 

 Education is simply too complex to permit the kind of certainty that characterizes the natural sciences, where a finding is a finding is a finding, where whatever was found to be true in Rio de Janeiro can be transferred to Los Angeles, or rural Mississippi, and on rainy as well as sunny days.

 

 Context matters in the social sciences. The context of a study is all of the circumstances that surround the putative causes and effects that the researcher is attempting to study: the locale, the time of year, the socio-economic level of the persons participating in the study. Each of these features of “context” may interact with the relationship of the independent and dependent variables – the cause and the effect – and change the nature of the relationship. Because of their complexity, we may never understand all the interacting influences that make up a particular context, and thus we may never be able to predict when and where a program will and will not work. But it’s more than the complexities of context that limit our confidence in a program’s transferability to a different setting. Three additional problems make it difficult to transfer programs that appear to work to a new and different setting.

 

 The Problem with Findings. First is the problem of estimating the power of the program that we want to import to our school or district. How strong were the original findings? Were the effects strong enough to suggest that we ought to try it elsewhere? Many reports of a successful program or activity present their results as “statistically significant.” But that doesn’t mean much because statistical significance is primarily a reflection of sample size. A pill that works for only one person out of 50 can produce a statistically significant result in a huge clinical trial. Interpreting data also requires knowledge of whether random assignment occurred and whether the investigators were the same people who developed the program under study. It is better to have data about a program’s effects presented as an effect size, which helps us decide whether the program’s effect, despite all the complications in the study’s design, is potentially large enough to be worth pursuing in terms of time, money, and personnel costs. 

 

 But even if the overall effect of a program was impressive, the conditions under which the program did not work are rarely discussed and are not well understood. The famous Tennessee class-size study (Mosteller, 1995), the STAR study, showed impressive overall benefits of smaller classes. Since that study was published, many have argued that major reductions in class size for poor children are likely to have lasting effects on the children’s lives.  But Konstantopoulos (2011) looked within the overall data and noted that results revealed that a large proportion of the school-specific small class effects are positive, while a smaller proportion of the estimates are negative. Although students benefit considerably from being in small classes in many schools, in other schools being in small classes is either not beneficial or is a disadvantage. Small class effects were inconsistent and varied significantly across schools in all grades. (p. 71) 

 

 This is no different a result from what we find in pharmacological studies. A drug may turn out to have an overall average positive effect, and thus is approved by the Food and Drug Administration. Forgotten in the rush to bring the drug to market are the data that show it didn’t work for many in the sample, it harmed some, and among those who showed positive effects were many people who responded because of placebo effects. Pharmacological research is closer to education research than research in the natural sciences is. 

 

  Just as human biological systems vary, and drugs work with some patients and not with others, school and class contexts vary a great deal. Programs like class-size reduction are fine candidates for improving the progress of poor students and the working conditions of teachers, but they may not always work as we hope. Konstantopoulos’s insights into the effects of the class-size study are similar to the advertisements for medicines one hears on television. You hear about how wonderful a drug is—just before the fast talk begins informing you that it may produce blood clots, susceptibility to tuberculosis, increased heart problems, and the like. We eventually learn that overall success is invariably accompanied by many noneffects and quite a few failures. 

 

  But few researchers, and even fewer promoters of programs, do the high quality research that would reveal noneffects, or negative effects for some children, when a given program is in the hands of some teachers and in certain schools. Education research doesn’t provide us with such answers.

 

 The Problem with Replicability. The gold standard of research is often said to be the randomized clinical trial. But we don’t think so. The real standard is a replication of effects by authors who neither produced the original study nor designed the original program. 

 

In medicine, one major study suggested that only 44 percent of the replications of medical research produced supportive data (Makel & Plucker, 2014). Unsuccessful replications most often occurred when the sample size in the original study was small and when randomization was not employed. These are precisely the conditions that describe a great deal of education research. But we don’t have a nonconfirmation problem in education research, as does medicine, because we have an even more serious problem: We don’t even do replication research!  The replication rate for research in our top journals, at well under 1 percent, is frighteningly low. The lack of replications, of course, makes it harder to be confident that a program that works in one location will work in another. 

 

The Problem with Fading Effects.   As teachers change, as student characteristics change, as assessment instruments change, and as school cultures change, a program that seemed successful a few years back may no longer work as it did. Programs need to be monitored for efficacy over time, just as medicines do. Also, ideas that are key to the program of interest may already be in place among the students we want to help, and so bringing the new program in shows little or no effect. 

 

 Lemons, Fuchs, Gilbert, and Fuchs (2014) examined five randomized studies of a supplemental peer-mediated kindergarten reading program involving more than 2,500 students across nine years. They found a dramatic increase in the performance of the control-group students over time. Obviously, if the control groups are doing better on the measures used to evaluate a program’s efficacy, it’s harder for the program to show an effect in a new district or school. The students in the control groups somehow were getting better instruction over time, so the power of the peer-mediated reading program to show its effects got weaker and weaker. We rarely have nuanced or complete data about the students we want to help when we bring in a new program, and this lack of understanding may weaken the effects we finally see. 

 

 The whole idea of “bringing programs to scale” (that is, moving a program from a few schools to many) is also a problem. Control of the contextual complexity in a few classes, or in a school or two, is a lot easier than control of the myriad contextual variables affecting programs in entire districts or states.  

 

 Realistically Optimistic

 

 So things don’t always work as expected. What are school leaders to do? The best they can! Some data are probably better than no data, if collected honestly by individuals who aren’t out to make a lot of money by pushing a program. 

 

 So look at the data. But overselling an idea or program in your own district is a mistake. You’ll need to try it out, probably adapt it to local circumstances, and then it still may not work as intended. But it might. A realistic view of the difficulties that lie in the path to school improvement must not lead to despair. As professionals, we’re expected to seek better ways of educating children. Trying out programs that have been successful elsewhere, designing new programs that fit local circumstances, and attempting to implement what sound like good ideas are characteristic of exemplary leadership. 

 

Three considerations will increase the chances that experimentation will lead to improvement. One is having teacher buy-in. Not much works well if teachers have things imposed on them that they don’t believe in. Second, don’t implement several new programs and ideas simultaneously. Teachers often suffer from overload when new administrators, or state and federal bureaucrats, set out to change too many things too quickly. Finally, make sure new programs and ideas undergo a formative evaluation to find out how things work and how they might be improved. This might entail asking a local evaluator or colleagues from a different school to help with formative and summative assessments of a program. 

 

In 1987, at the signing of a treaty with the Soviet Union, President Reagan remarked, “Trust, but verify.” His advice is our advice: Trust that your colleagues across the United States and around the world have found some good ideas for school improvement that work for them. But verify that their thinking will work for you, too. EL  

 

Postscript

 

Ideas That (May) Travel Well

 

Here are a few pet ideas that we’ve seen work in one place or another that might offer alternative approaches to school improvement:

 

 * Stop looking for answers to local problems in Scandinavia or Asia. The United States is neither Finland nor Singapore, and it’s a lot more complex than either.

 

 * Redraw school attendance areas to achieve socioeconomic balance, and support high-quality early childhood education in those areas.

 

 * Recognize that teachers work in teams and evaluate them accordingly. Make sure the evaluation system has no consequences for teachers associated with student test scores but does include multiple classroom observations and an evaluation of classroom artifacts—tests, papers, projects, and the like.

 

 * Eliminate tracking in grades K–6, and eliminate grade retention (“flunking”) completely.

 

 * Make sure that no school day for students starts earlier than 8:30 a.m.

 

 * Provide libraries staffed with librarians and counseling offices staffed with enough counselors that they can know students personally.

 

 * If you don’t like your reading scores, find ways to have students read more, and forget most other systems that claim to improve reading. There is no "Science of Reading."

 

   References  

 

 Berliner, D. C., Glass, G. V, & Associates. (2014). 50 myths and lies that threaten America’s public schools. New York: Teachers College Press. 

 

 Fernandez, M., & Schwartz, J. (2013, December 13). Teenager’s sentence in fatal drunken-driving case stirs “affluenza” debate. New York Times. Retrieved from www.nytimes .com/2013/12/14/us/teenagers-sentencein- fatal-drunken-driving-case-stirs-affluenza- debate.html

 

 Konstantopoulos, S. (2011). How consistent are class size effects? Evaluation Review, 35(1), 71–92.

 

  Lemons, C. J., Fuchs, D., Gilbert, J. K., & Fuchs, L. S. (2014). Evidence-based practices in a changing world: Reconsidering the counterfactual in education research. Educational Researcher, 43(5), 242–252. 

 

 Makel, M. C., &. Plucker, J. A. (2014). Facts are more important than novelty: Replication in the education sciences. Educational Researcher, 43(6), 304–316.

 

Mosteller, F. (1995). The Tennessee study of class size in the early school grades. Future of Children, 5(2), 113–127.  

 

 

 

Friday, May 10, 2024

Meta-analysis of the effects of type and combination of feedback on children's discrimination learning

1985

Getsie, R.; Langer, P.; & Glass, G.V (1985). Meta-analysis of the effects of type and combination of feedback on children's discrimination learning. Review of Educational Research, 55(1), 9-22.

Saturday, May 4, 2024

Review of Smith and Adams's Educational Measurement for the Classroom Teacher

1967

Glass, G.V (1967). Review of Smith and Adams's Educational Measurement for the Classroom Teacher. Educational Forum, 31, 245-246.

David Berliner's Legacy-Loss of a Friend: David Charles Berliner (1938-2025)

David Berliner's Legacy-Loss of a Friend: David Charles Berliner (1938-2025) Gene V Glass Professor Emeritus Arizo...