IEEE TRANSACTIONS ON VISUALIZATION AND COMPUTER GRAPHICS 1077-2626/98/$10.00 © 1998 IEEE
Vol. 4, No. 4: OCTOBER-DECEMBER 1998, pp. 307-316
Perception of Human Motion With Different Geometric Models
*![]()
Jessica K. Hodgins, Member, IEEE
James F. O'Brien, Student Member, IEEE
Jack Tumblin
Abstract Human figures have been animated using a variety of geometric models, including stick figures, polygonal models, and NURBS-based models with muscles, flexible skin, or clothing. This paper reports on experimental results indicating that a viewer's perception of motion characteristics is affected by the geometric model used for rendering. Subjects were shown a series of paired motion sequences and asked if the two motions in each pair were the same or different The motion sequences in each pair were rendered using the same geometric model. For the three types of motion variation tested, sensitivity scores indicate that subjects were better able to observe changes with the polygonal model than they were with the stick figure model.
Index Terms Motion perception, motion sensitivity, computer animation, geometric model, perceptual study, biological motion stimuli, light-dot display.
1. INTRODUCTION Few movements are as familiar and recognizable as human walking and running. Almost any collection of dots, lines, or shapes attached to an unseen walking figure is quickly identified and understood as human. Studies in human perception have displayed walking motion using only dots of light located at the joints and have found test subjects quite adept at assessing the nature of the underlying motion [1]. In particular, subjects can identify the gender of a walker and recognize specific individuals from light-dot displays, even when no other cues are available [2], [3], [4].
In part because people are skilled at detecting subtleties in human motion, the animation of human figures has long been regarded as an important, but difficult, problem in computer animation. Recent publications have presented a variety of techniques for creating animations of human motion. Promising approaches include techniques for manipulating keyframed or motion capture data [5], [6], [7], [8], control systems for dynamic simulations [9], [10], [11], [12], [13], and other procedural or hybrid approaches [14], [15], [16], [17], [18], [19], [20]. Each method has its own strengths and weaknesses, making the visual comparison of results essential, especially for the evaluation of such subjective qualities as naturalness and emotional expression. The research community has not yet adopted a standard set of models and there is currently enormous variety in the models and rendering styles used to present results.
Our ability to make judgments about human motion from displays as rudimentary as dot patterns raises an important question: Does the geometric model used to render an animation affect a viewer's judgment of the motion or can a viewer make accurate judgments independent of the geometric model? There are three plausible but contradictory answers to this question.
Possibility 1. Simple representations may allow finer distinctions when judging human motion. Simpler models may be easier to comprehend than more complex ones, allowing the viewer's attention to focus more completely on the details of the movement rather than on the details of the model. For example, a stick figure is an obvious abstraction and rendering flaws may be easily ignored. When more detailed models are used, subtle flaws in rendering, body shape, posture, or expression may draw attention away from the movements themselves. Complex models may also obscure the motion. For example, the movement of a jacket sleeve might hide subtle changes to the motion of the arm underneath.
Possibility 2. Complex, accurate representations may allow finer distinctions. People have far more experience judging the position and movement of actual human shapes than they do judging abstract representations such as stick figures. A viewer, therefore, may be able to make finer distinctions when assessing the motion of more human-like representations. Furthermore, complex representations provide more features to identify and track. Each body segment in a polygonal human model has a distinctive, familiar shape, making it easier to gauge fine variations in both position and rotation.
Possibility 3. Both simple and complex representations may allow equally fine distinctions. The human visual system may use a displayed image only to maintain the positions of a three-dimensional mental representation. Judgments about the motion may be made from this mental representation rather than directly from the viewed image. Displayed images must, of course, supply enough cues to keep the mental representation accurate, but additional detail and accuracy may be irrelevant. Just as joint positions shown by light dots are sufficient to control the mental representation, connecting the dots with a stick figure might not improve the viewer's perception. Similarly, encasing a stick figure within a detailed human body shape might likewise prove unnecessary.
Objective evidence is needed to determine which of these possibilities is correct. We argue that definitive experiments to select between Possibilities 1 and 2 are impractical. The question of which style of geometric model is more useful for judging motion is likely to be highly complex and context dependent, affected by all of the variables of both the motion and the rendering. If Possibility 3 were correct and model style were largely irrelevant, then we would be able to perform critical comparisons of the motion synthesis techniques in the literature by direct comparison of the substantially different geometric models used in each publication. This paper provides experimental evidence to disprove Possibility 3 by showing that viewer sensitivities to variations in motion are significantly different for the stick figure model and the polygonal model shown in Fig. 1. In particular, for the types of motion variation we tested, viewers were more sensitive to motion changes displayed through the polygonal model than through the stick figure model. This result suggests that stick figures may not always have the required complexity to ensure that the subtleties of the motion are apparent to the viewer.
Fig. 1. Images of an animated human runner. (a) Two running motions rendered using a polygonal model. (b) The same pair of motions are rendered with a stick figure model. Modifications to the motion were controlled by a normalized parameter,
, that varied between
= 0 and
= 1. These images are from the motion generated for the additive noise test discussed in Section 3.3. The difference in posture created by the additive noise can be seen in the increased angle of the neck and waist in the right image of each pair (
= 1).
2. BACKGROUND Several researchers have used light-dot displays, also referred to as biological motion stimuli, to study perception of human movements and to investigate the possibility of dynamic mental models [21]. The light-dot displays show only dots or patches of light that move with the main joints of walking figures (Fig. 2), but even these minimal cues have been shown to be sufficient for viewers to make detailed assessments of the nature of both the motion and the underlying figure.
Fig. 2. (a) The joint locations of a human runner at a single point in time. (b) The joint locations over the course of one step in the running cycle. Although it is difficult to determine the nature of these patterns from a still image, studies show that most people are able to recognize the motion and even to make fine judgments when shown moving sequences of similar images.
The ability to perceive human gaits from light-dot displays has been widely reported to be acute and robust. Early experiments by Johansson reported that 10-12 light dots "evoke a compelling impression of human walking, running, dancing, etc." [1]. Because such displays provide motion cues independent of form or outline, other investigators have used them to study human motion perception. Work by Cutting and Kozlowski showed that viewers easily recognized friends by their walking gaits on light-dot displays [2]. They also reported that the gender of unfamiliar walkers was readily identifiable, even after the number of lights had been reduced to just two located on the ankles [3]. In a published note, they later explained that the two light-dot decisions were probably attributable to stride length [4]. Continuing this work, Barclay et al. showed that gender recognition based on walking gait required between 1.6 and 2.7 seconds of display, or about two step cycles [22]. Our experiments used pairs of running stimuli 4 seconds in duration that displayed about six strides. We noticed that test subjects often marked their answer sheets near the midpoint of the second stimuli, which is consistent with Barclay's results.
Motion is apparently essential for identifying human figures on light-dot displays. The Cutting studies reported that, while moving light-dot displays were recognized immediately, still light-dot displays of a walking figure were not recognized as human. Poizner et al. also noted that movement is required for accurately reading American Sign Language gestures [