Sunday, December 19, 2010

Tales of the Rampant Coyote: Exception Driven Game Play

Rampant Coyote has an excellent post on exception driven game play in RPGs which ties in rather nicely with my series on randomness in RPGs(and not for the first time). It's a must read if you often find modern PC RPGs lacking compared to their less refined ancestors.

Labels: , , ,

Thursday, May 21, 2009

Tales of the Rampant Coyote: RPG Design: In Defense of ... Hit Points

Another great post at Tales of the Rampant Coyote. This one discusses the pros and cons of "hit points" in RPG design. The comments are especially interesting and include discussion of pen and paper classics like Rolemaster and Cyberpunk as well as new electronic games like Dwarf Fortress.

Labels: , ,

Friday, May 08, 2009

Tales of the Rampant Coyote: RPG Design: That Which Is Not Forbidden...

Very interesting post at Tales of the Rampant Coyote which ties in directly with my concept of player input variance and open-endedness in RPGs. In discussing two of Gary Gygax's published pen-and-paper adventures:
But both Tomb of Horrors and Necropolis left a lot up to interpretation by the Dungeon Master (the person who "runs" the game). And I try and run my games by a guiding rule which, lamentably, tends to be ignored in more recent editions of the game, and ignored by players who are used to computer games: That which is not expressly forbidden is fair game to try.
And the conclusion:
The problem is that - for the most part - RPGs aren't made as anything resembling simulations. That's too difficult, and it is too hard to put the player on the kinds of rails that many designers prefer. So spells have very particular, extremely limited uses, and tend to be more of the "blow crap up" variety. Spells that provide knowledge, hints, or "intelligence" are subject to exploit in single-player games, as the information they provide to the player is persistent, even when the player reloads the game immediately to 'restore' the expended spell.

Our worlds are just too restrictive to allow this kind of play. But do they have to be?
They definitely do not have to be so restrictive. One way to deal with the issue is through random procedural content. Another is by limiting the ability to save/reload. Roguelike developers already care about these issues. I'm just waiting for the mainstream to catch on and it makes me very happy to see someone else talking about it.

Labels: , , ,

Thursday, May 07, 2009

Randomness and the Black Swan in RPGS(Part 2)

In the last two entries of this series I discussed the role of randomness in RPG task resolution mechanics. I also demonstrated the current shift toward increasingly deterministic resolution systems in electronic RPGs, and I explained why I felt this was a bad thing. Now it is time to don the mantle of chaos once again and discuss the other major set of RPG systems influenced by randomness, content generation systems. I will also introduce a couple of conceptual tools which may prove useful in my analysis.

In most P&P RPGs, randomness exerts its influence from the very beginning, at character creation. Anyone who has played any edition of Dungeons and Dragons is familiar with summing the results of 3 6-sided dice to determine a character's ability scores like strength, intelligence, etc. Hit points for both players and monsters are also determined by rolling dice. These are all simple ways in which randomness feeds into the content generation systems of RPGs. Simple randomness ensures that all orcs aren't exactly the same. The inclusion of this type of randomness also means that one player's 3rd level fighter may be more powerful than another player's. However, the greater the number of attributes with randomly determined parameters, the more likely most characters will trend towards average overall and the more likely most characters will excel in some area.

Random character generation was one of the first random elements of P&P RPGs to be removed from electronic RPGs. Aside from licensed D&D games and roguelikes, few electronic RPGs still have random elements in their character generation systems. In many electronic RPGs there isn't even such a thing as character generation. This is entirely logical for story-based games and games without much in the way of randomized content generation. But, when the challenges to be met aren't predetermined, why shouldn't an RPG utilize randomness from the very beginning to vary the game experience?

But when most people think of random content generation, they think of randomly generated environments, encounters, and rewards. All of these elements were present in D&D in the 1970s. In general, every encounter area was expected to have an associated wandering monster table. The DM made periodic wandering monster checks, and if one occurred, the appropriate table was consulted. This allowed the DM to generate encounters on the fly while the party was travelling long distances or camping. Wandering monster checks could also be used to punish players for making foolish decisions like making a lot of noise in a dungeon. The AD&D Dungeon Masters' Guide included encounter tables for all types of wilderness terrain and dungeon levels. There were also tables to generate treasure on the fly for these and other encounters. Each monster type had an associated treasure type with an associated set of tables for randomly determining what treasure a monster of that type might have in its possession. There was a very slight chance that even a weak creature could have a powerful magical item.

With the release of the AD&D Dungeon Masters' Guide in 1978, random content generation was taken to the next level with the inclusion of a complete system for randomly generating dungeons. Level layouts, environmental details, tricks, traps, treasure, and enemy encounters could all be generated on the fly by the DM. The inclusion of such a system says volumes about the way at least some players played P&P RPGs. It is also some indication that, at least for some players, the mechanics of D&D were fun in and of themselves.

Of course, it should be mentioned that the Dungeon Master's Guide explicitly states that the content generation charts contained therein are not to be taken as Law. The DM can freely ignore any result he feels will unbalance the game or which he simply doesn't like. However, potentially unbalancing random events are ameliorated somewhat by the flexibility of a game played between humans. If the players suffer a stroke of bad luck and encounter a high level wandering monster they have a practically infinite number of options available to them, limited only by their creativity. They can scatter in all directions, they can describe in detail their attempts to hide or escape, etc. The DM can then reward quick thinking and good ideas. Contrast this with early electronic RPGS where the only option besides fighting was selecting RUN from a battle menu and it is easy to see why the range of random encounters might have been restricted.

Now I would like to introduce two concepts: content generation variance and player input variance, hereafter referred to as CGV and PIV. CGV is a measure for the range of output possible by the content generation system given a fixed set of parameters. CGV is easy to quantify in a video game because it is essentially an algorithmic function with specified parameters. PIV is harder to quantify because it always involves human input. Sometimes, such as with menu systems, the human input is very restricted, but many modern RPGs are much more open ended than that. In addition, PIV is also a result of the way in which human input is processed and translated into game actions. PIV takes into account randomness within the game's task resolution systems(see Part 1 for more details), for instance. So, PIV is defined as the range of results from all possible player actions in a given game situation. PIV will remain a fuzzy concept, but will suffice for our purposes. A P&P RPG played between humans will be considered as having maximum values for both CGV and PIV.

Next time I will continue the discussion of CGV and PIV. I will discuss how they relate to one another and how they have changed over the years in RPGs.

Labels: , ,

Friday, February 13, 2009

Randomness and the Black Swan in RPGS(Part 1.5)

Before looking at the random generation of game content in Part 2, I want to elaborate on a point made in Part 1 concerning combat mechanics: "The point of this comparison is that predictability isn't necessarily more fun because it is "balanced", sometimes it is just boring. And if it would be boring in a pen and paper game, then fundamentally, it is probably still boring electronically. It is just dressed up in a such a way that people put up with it."

Admittedly, I glossed over all the elements which can compensate for predictability in a electronic RPG's combat system. For one, the game can require the player to manage an assortment of abilities/skills in a real-time situation where tactical decisions must be made quickly and the player has to adapt to evolving situations. This is the approach taken by most MMORPGs, and some, like World of Warcraft, do it well. I still argue that too much predictability is a negative, though. Few encounters in such a system are unique and there is a very clear delineation between encounters the player can handle and encounters he can't. It is common for a player to find himself in a situation where if he draws/lures one enemy he will win fairly easily, but if he draws two he will die. This type of certainty reduces the depth of decision making. The challenge is moved firmly into the realm of player execution.

Interestingly, Final Fantasy 12, which I used as an example of predictable combat, seems to recognize that its base combat mechanics aren't very interesting on their own. Recognizing that the opportunities for complex decision-making in Final Fantasy combat are fairly rare, the developers automated the system. The combat in Final Fantasy 12 plays out in real-time according to simple, player-assigned AI rules controlling each character's actions. The player can pause the action and intervene if necessary, but this is rarely required outside of boss battles. Instead, the player takes satisfaction from tweaking his party's AI until it is as self-sufficient as possible.

The lead developer of Battle for Wesnoth, an open source turn-based strategy game, explains why his game has a substantial random element in its combat mechanics here. He makes a number of excellent points that apply just as well to role-playing games:

In Wesnoth we want a player to plan out a complex situation, to estimate carefully the possibilities. To have to work out a good strategy. If a player can simply rely on all sorts of assurances that their units will hit, they don't have to do any of this. Sure, there will be a certain amount of fun to planning out a situation where you can set up a cool 'domino effect' of enemy units going down as you attack them. But this is nothing to do with the skill of planning out a real strategy in a dynamic situation where you have to consider all kinds of contingencies.

...

Trying to "slide down the scale" of luck would simply make Wesnoth less interesting, in my view. Suddenly the difference between a two attack unit and a four attack unit would be trifling instead of critical. Not near as much planning would be necessary. Simply, in my view, Wesnoth would become less fun.


The point I want to make in this post is that although predictability can be compensated for, especially if the combat system plays out in real-time, an infusion of unpredictability into the system can make it even more interesting.

Labels: , , ,

Wednesday, February 04, 2009

Randomness and the Black Swan in RPGs (Part 1)

Raph Koster wrote an interesting post on the connection between games and the ludic fallacy as discussed in Nassim Nicholas Taleb's bestselling book The Black Swan. For those unfamiliar with the book, a black swan is an unpredictable but highly impactful event. Raph ends his post with a question:

"Rather than just bemoan this, I'll instead issue the challenge: what is the fun game that features black swans, phase transitions, and the catastrophic 100-year flood? How do you sculpt a system that does this without chasing away the newbies?"

This is an interesting question, and one which has a simple answer if not constrained to electronic games: pen and paper roleplaying games. The presence of a human game master lends unpredictability, surprises, and ad hoc creativity that electronic games can't hope to match. However, even ignoring the human element, the mechanics of electronic RPGs have been straying ever further away from unpredictability in a misguided pursuit of balance. Since long before I had the Black Swan framework to hang it upon, I have been thinking about the role of randomness in roleplaying games.

Roleplaying games typically draw on randomness for two different major functions. The first is in their action resolution mechanics. The second is in their content generation systems. Both of these got their start in the earliest pen and paper roleplaying games. The earliest electronic RPGs, existing primarily to recreate the tabletop experience, inherited both of these random elements. As electronic RPGs have diverged from their pen and paper roots, the role of randomness has steadily diminished. The results of this process are manifested most clearly in today's MMORPGs and JRPGs. But let's not get ahead of ourselves, even within the early pen and paper tradition, randomness has been afforded varying degrees of importance.

Let us begin by comparing the combat resolution systems of Advanced Dungeons and Dragons and Rolemaster. In AD&D a typical orc has 1d8 hit points and an armor class of 6. A first level fighter with no bonuses will hit an armor class of 6 with a roll of 14 or higher on 1d20(the standard die used for attack rolls in AD&D). At tenth level, the same fighter only needs to roll a 5 or higher to hit the orc. With a standard longsword, the fighter will do 1d8 hit points of damage to the orc if his attack succeeds. So, the first level fighter will hit the orc 35% of the time and it will take anywhere from one to eight successful blows to kill the orc. According to my calculations, the chance the fighter will kill the orc with his first swing is just under twenty percent(0.35 * 0.5625). However, if the fighter had a strength of 18/00(the max a human fighter can have naturally), he would receive a +3 bonus to hit and a +6 bonus to his damage roll. In this case, the fighter would have almost a fifty percent chance to kill the orc on his first attack. Clearly, there is a large random element to combat in AD&D. Bonuses from high ability scores and special abilities and equipment do reduce the relative importance of the random rolls, but these bonuses are usually fairly small. The main constraint to AD&D randomness is that the range of results is clearly defined. If a player's character has 10 hit points and an enemy is attacking him with a longsword, the player knows that his character can't be killed in one blow unless the opponent has some damage bonuses from strength, magic, etc.

In Rolemaster, however, things aren't quite so simple. The attack roll in Rolemaster is an open-ended d100. Normally, results range from 1-100, but on a natural/unmodified roll of 96-100 another d100 roll is made and added to the result of the first roll. This second roll is open-ended as well, so a natural 96-100 will result in a third roll, and so on, ad infinitum. The upper bound of an attack roll in Rolemaster is infinity. Unlike AD&D, the attack roll determines both whether a hit occurred and how much damage was inflicted, with one exception. The final attack roll result, plus the attacker's offensive bonus and minus the defender's defensive bonus, is indexed on a chart for the weapon being used and cross-referenced against the defender's armor type. The result is a number indicating the amount of damage inflicted and possibly a letter indicating a critical hit. Most combat deaths in Rolemaster are either direct or indirect results of critical hits; combatants tend to die in sudden and dramatic fashion and not as much through attrition. There are five degrees of critical hits, A through E, with E being the deadliest. When a critical is indicated, a separate d100 roll is made on the appropriate critical chart. Crticals can result in broken bones, bleeding, blindness, loss of consciousness, and even gruesome death. Even an A critical can kill on a natural roll of 100. Combine the lethality of critical hits with the fact that even the best armor usually can't reduce the probability of a critical hit to zero, and you have a system where any blow can kill. The genius of Rolemaster is that it allows players to manage risk by using any portion of their offensive bonus as a defensive bonus instead. It is rare for players to take a purely offensive stance unless their opponent is stunned.

Now let us draw a comparison between these older pen and paper mechanics and a modern console RPG, Final Fantasy 12. The damage equation for wielding a spear is:

DMG = [ATK x RANDOM(1~1.125) - DEF] x [1 + STR x (Lv+STR)/256]

Notice that the random element ranges from 1.0 to 1.125. Vaan at 50th level, wielding the best weapon in the game(ATK 150) against an enemy with a DEF of 50 yields:

DMG = [150 x RANDOM(1~1.125) - 50] x [1 + 50 x (50+50)/256]

The end result is damage ranging from 2053 to 2438, a fairly narrow range compared to the pen and paper systems. Also notice that successful attacks are the default. Instead, certain equipment like shields provide a chance block an enemy's attack. There is also a small chance to score a critical hit, which results in double damage. As you can clearly see, combat resolution in Final Fantasy 12 is much less random than even AD&D.

To see for yourself how much things have changed since the early days of roleplaying, just compare the combat of a grind heavy MMORPG with the commercial MUD Gemstone IV. Gemstone IV's game mechanics are derived from Rolemaster(in fact, earlier versions of the game licensed the Rolemaster rules and the Shadow World setting). You will not hear Gemstone players using terms like DPS(damage per second) when discussing combat strategy. You often hear complaints from MMORPG players concerning the grind necessary to gain levels, but in fact every MMORPG battle is a microcosm of the level grind. This is especially true of the large raid bosses. You can find a very brief excerpt of a Gemstone IV battle here. I was caught in an invasion of high level plant creatures into my low-level hunting ground and killed. Another player who attempted to rescue me was killed as well. The point of this comparison is that predictability isn't necessarily more fun because it is "balanced", sometimes it is just boring. And if it would be boring in a pen and paper game, then fundamentally, it is probably still boring electronically. It is just dressed up in a such a way that people put up with it.

But even the relatively unpredictable and complex combat mechanics of Rolemaster fall short of the conditions needed for Taleb's black swans. Even without knowing all the bonuses involved, the player can roughly estimate the risks. Looking at the attack charts, the player can see what results are needed to generate critical hits. With transparent mechanics, black swans are by definition impossible because the possibilities are known beforehand. With regard to task resolution mechanics, this isn't necessarily a bad thing. It is the generation of game content which can really benefit from black swans -- like the plant invasion which killed my character in Gemstone IV -- and this shall be the topic of my next post in the series.

Labels: , , ,

Monday, June 16, 2008

On Tetris

A commenter to one of my earlier L/N posts asked what N-Score I would give Tetris. It is an interesting question and one I considered myself when I first began thinking about L/N.

To begin with, the N-Score would depend on the specific version of Tetris being reviewed. Any version of Tetris with a decent soundtrack and nice, crunchy sound effects would score somewhere in the average range of 9 to 12. A lack of, or low quality, sound effects or music could bring the N-Score below average while a truly excellent soundscape could pull it a little above average. Remember, I score games relative to their concept, so a complex narrative isn't a requirement for a puzzle game like Tetris. However, there are a number of things that a Tetris game could do to earn a significantly better than average N-Score.

For one, a dynamic soundtrack could do a lot to improve the "narrative experience" of Tetris. If the music matched the state of your current game, whether frantic and tense as you play on the knife's edge of defeat or triumphal and celebratory after landing a well set up tetris, you would certainly be more immersed in the game.

A Tetris game could also take a few cues from mahjong games and add opponents or antagonists that you play against. Instead of pieces dropping without explanation from the top of the screen, they could be positioned and dropped by an animated character. Tougher opponents would drop them faster and with different algorithms.

A Tetris game which did both of these could potentially earn a very high N-Score. So really, high N-Scores are not limited to certain narrative centric genres. Any game which maximizes the artistic and narrative potential of its concept can earn a high N-Score.

Labels: , , ,

Sunday, June 15, 2008

L/N Implementation Details

In an earlier post I described a dual metric for reviewing video games called the L/N system. That first post details both my philosophy of assigning quantitative scores and the qualities to be measured by each metric. Now it is finally time to discuss the specifics of my implementation of the L/N system. I say 'my implementation' because the L/N concept itself only applies to the what and doesn't extend to the how. Other reviewers are encouraged to develop their own implementations of the system.

A Problem of Scale

Probably the most common set of complaints I see regarding quantitative scores relate to the scale being used, which is almost always linear. Sometimes there is confusion in interpretation because the scale used is numerically linear but semantically nonlinear. For instance, the difference between a 6 and a 7 is greater or less than the difference between a 9 and a 10. Another common source of misinterpretation regards what constitutes an average score. Percentile scales and 10 point scales are particularly susceptible to this type of misinterpretation. Many readers will interpret a 7 or 70% as an average score even though it is well above the median score of 5 or 50%. The print magazine EGM recently moved from a 10 point scale to a letter grade scale to mitigate this sort of ambiguity. To avoid confusion, the scale used should be numerically and semantically consonant. The numerical mean should match the semantic mean and the relative numerical differences should be consistent with their semantic interpretation.

Another common complaint related to numerical scales is that the differences between scores often seem arbitrary. Percentile scales, for instance, are much more precise than any reviewer can justify. Having too much precision in a scale will undermine it to the point that readers will begin to question the accuracy as well. Linear scales compound the precision problem because the greatest works are often much, much more impressive than the average, but extending a linear scale far enough to adequately represent that may make the scale appear to be overly precise.

My Solution

So, I want to avoid linear scales and use a nonlinear one that gamers have an intuitive grasp of. It is also best if the scale is not easily confused with a linear one. My solution is to use a scale based on 3d6. Yep, the very same scale used to measure ability scores in Dungeons and Dragons(for those not in the know, D&D players generate their characters' ability scores by rolling three 6-side dice and summing the results). The biggest numerical advantage to using 3d6 is that it is distributed normally, i.e. the graph of possible values is a bell curve. The major non numerical advantage is that it is familiar, and even if you don't know what a normal distribution is, you probably realize that rolling an 18 is quite a bit harder than rolling a 17. The scores themselves will be represented visually by images of three dice. This serves to add another layer of information based on the specific dice chosen to represent the score. A 13 could be represented by 5-4-4 or 6-6-1, for example. A score of 6-6-1 would indicate potential greatness brought down by one or more serious flaws.



To provide a little more detail, the median score is 10.5. Since I'm obviously only using whole numbers, this means the average score ranges from 9 to 12. Scores within this range account for roughly 48% of the population. Less than half of one percent of the population would have an 18. Of course I am not going to assign scores with the sole aim of fitting a probability distribution, but the distribution does help define the relative difference between scores.

Guidelines For Assigning Scores

Now that I have defined the scale, I'd like to put forth a few guidelines I will follow when assigning scores. A number of questions have been raised by readers so far and hopefully these guidelines will paint a clearer picture of my scoring criteria.

First, context is important when scoring a game. The very highest scores are reserved for those games which are both incredibly well executed and groundbreaking in some way, and innovation is meaningless outside of the context of when the game was released. The original DOOM would score higher than the many clones which followed, even though some may have been just as well executed. Technological context is also important. Were a new 16 bit console game released today, its technology-related aspects would not be compared against current gen consoles.

In addition to historical and technological context, the concept behind the game is also important. Many reviewers already rate games relative to their genre, but I would like to be explicit in my belief that reviewers should take it a step further and make an effort to divine the developers' specific ideals and goals. I am a firm believer that not all games, not matter how excellent, will appeal to all players. If the very concept of a game and what it is trying to achieve doesn't appeal to a player, then they probably will not like it. A game should be judged in light of these factors. However, this doesn't mean that the concept itself is beyond reproach. Lack of ambition, for example, will certainly prevent achieving the highest scores.

For determining the N-Score, a useful guideline is to consider how interesting the game would be if you were watching someone else play it. This doesn't perfectly capture my concept of the N-Score, but it is still useful to consider because many elements measured by the N-Score can be appreciated by someone other than the player, such as music, art design, story, etc. There are some crucial differences though. Player immersion is something that the N-Score should be concerned with, but immersion is hard to measure if you aren't the one playing. The visual feedback provided by video games in response to user input can make the user feel like a hero in ways that movies and other passive media cannot. This type of immersion isn't captured by the L-Score because the visual feedback often has little or no impact on actually playing the game. It is a complex topic, but hopefully these examples clarify the concept of the N-Score.


And with that, I believe I have written more about my video game review philosophy and metrics than any mainstream gaming web site or print mag. Kind of a shame considering that I haven't written a single review and they've written thousands.

EDIT: Added a paragraph on 'concept' to the guidelines.

EDIT: Eurogamer actually has a fairly detailed description of their scoring policy. It doesn't address the same problems as my system, but at least it provides a detailed semantic description of their scale.

Labels: , , ,

Monday, June 09, 2008

A Defense of Game Review Metrics

The subject of video game reviews turned out to be a hot topic last week. In addition to my post Reviewing and Scoring Video Games, there was a column on gamesetwatch by Simon Parkin and an interesting article at PopMatters by L.B. Jeffries. Jeffries makes an excellent point in distinguishing the majority of game reviews today from the type of real criticism that the industry could use a lot more of -- reviews are targeted towards consumers making purchasing decisions while criticism is targeted towards the game makers themselves. According to Jeffries, most reviews today don't go beyond "this game is/isn't fun" to explore the why's which could help game developers make better games. Parkin also has a few points to make concerning consumer oriented reviews. Parkin contrasts video game reviews with consumer electronics reviews, noting that an objective measure of quality isn't possible for games in the way that it is for consumer electronics. Instead, Parkin says, game review scores are really a measure of how well a game lives up to its pre-release hype, although consumers still view them as an objective measure of quality.

In reading the comments to these articles and others, I've gathered that many people agree with Parkin's point that attempts to objectively rate games are fundamentally flawed. While I agree that pure objectivity is impossible, there are a number of reasons why quantitative metrics are still worthwhile:

1) Quantitative metrics allow for searching and sorting. If a reader finds a critic he often agrees with, he can quickly find all games that the critic rated highly without having to skim the text of every review.

2) Quantitative metrics allow for algorithmic processing and analysis. Even though "wisdom of crowds" aggregating sites such as metacritic are often flawed, one shouldn't condemn the entire concept. Metacritic has a lot of problems, but most are related to the site's implementation. A critic's review scores could even be used to rate the critic himself. The potential applications are endless.

3) Multidimensional metrics provide a framework for the reviewer, hopefully improving consistency when assigning scores. Flaws in games are naturally more apparent when the game is judged from different perspectives, reducing the likelihood of a reviewer reflexively handing out a perfect score to a flawed game merely because it does some things better than any game which came before it.

Futhermore, enough people like review scores to prevent them from going away anytime soon, so we might as well spend a little time thinking about creating better metrics.

Out of all the critics of game review metrics, the group I most respect are those, like Jeffries, calling for more insightful criticism and less consumer oriented reviews. I also consider this to be a very real problem, but it doesn't entirely preclude the use of metrics. Certainly there are many focused pieces of criticism which wouldn't have anything to gain by applying a numerical rating, but more macroscopic pieces which analyze the entire work could still gain a lot from quantitative metrics. I, for one, plan on writing reviews that utilize both metrics and, hopefully, insight.

Labels: , , ,

Sunday, June 01, 2008

Reviewing and Scoring Video Games

I've been considering writing a few game reviews for my blog, which inevitably leads to thinking about scoring systems. Assigning a concrete score to any creatively produced work isn't something to take lightly. If a grade is assigned, it naturally creates an aura of objectivity and carries the weight of perceived authority. In many cases, the grade assigned carries more weight than the content of the review itself. The final score also opens the critic to criticism as well. If the critic desires the air of authority that concrete scores engender, he must take as much responsibility for the score assigned as he does for the content of his review.

It is for all these reasons that a critic should think carefully about any scoring system that he adopts. It is vital that the system used is consistent with the critic's philosophy of judging the medium in question. For me, the act of assigning a score of some sort is important because I believe that works of art CAN be judged objectively. I wouldn't bother with criticism at all if I didn't feel that this was the case. The challenge is to devise a scoring system which is informative enough to allow readers with their own varying predilections to make their own interpretations of quality without sacrificing the objectivity and finality of assigning a 'final' score. It is difficult to do this with a single, one dimensional metric.

In the old days game magazines would rate games on graphics, sound, difficulty, etc. Breaking the score down to this level of granularity is problematic for a couple of reasons. First, I might not be an expert in every category I might determine is necessary to judge. I feel much more qualified to judge a game's graphical quality than I do its sound design, for example. Second, it is important for a critic to take a stand on excellence, to make a final judgment. A myriad of small judgments certainly doesn't carry the same weight as one definitive score. And finally, the metrics used to describe one work's greatness may not paint an accurate picture of another.

Speaking of video games specifically, academics in the field of game studies can be roughly divided into two different camps, the narrativists and the ludologists(the wikipedia entry for ludology has a brief description of the differences for the uninitiated). I have yet to see a game scoring metric which synthesizes the current academic discussion on games. Therefore, I am proposing the use of a system which consists of two scores, one measuring the game's excellence from a ludological perspective and the other rating the narrative as it applies to the game. For lack of better terminology I will refer to these as the L-Score and N-Score, respectively. I am personally more of a ludologist, but that doesn't obviate the importance of narrative elements. After all, people play games for different reasons.

The L-Score is the score which is most closely related to the uniqueness of the medium. I have argued before that games are different from art because they aren't simply admired, they are also played. It is the interactive nature of games which ludologists emphasize, and so one can think of the L-Score as a metric for gameplay and game design. Mechanics, systems, and level design are the key components measured by the L-Score.

If the L-Score is a measure of a game's design, then the N-Score is a measure of its artistic achievement. The narrative, in this case, is defined rather broadly. It consists of the game's music, writing, visual style, sound design, overall setting, etc. All of these factors influence the player's involvement in the game and are therefore important even if they don't have much of a direct impact on the actual gameplay.

There may be some overlap between the components measured by the L-Score and the N-Score. For instance, the sound design in a first person shooter may provide an increased level of information and awareness to the perceptive player. Such a feature could be considered relevant to both the N-Score and the L-Score. Likewise, in an exploration intensive RPG interesting environments may be necessary to realize the goals of the game's design, making those environments important from a design perspective as well as an artistic one. Despite any overlap between what is being measured by the two metrics, each metric is still able to stand on its own.

All games are scored relative to what they are trying to achieve, with the very highest scores reserved for true innovation. The traits that make a good RPG are simply quite different from those of an action game, and so the game's concept must of course be in mind when considering the quality of the game's design. Similarly, when judging a game's narrative it would be silly to expect the same level of exposition from a shmup as from an RPG. The narrative of a shmup is less about plot and more about evoking a certain feeling through music and visual presentation. Genres which are more narratively focused will in some ways be judged to a higher standard. The fact that many story-focused RPGs require 40+ hours to finish places a huge burden on developers to create a consistently strong narrative and interesting setting. A five stage shmup should not be punished for having less content(unless more content would make for a better shmup.)

There has been a lot of debate recently concerning game review scores, with several print magazines altering or eliminating their review scoring system(EGM and Play, respectively.) I believe the main reason for dissatisfaction with most current game review metrics is that they no longer accurately reflect gamers' increasingly sophisticated view of the medium. Games are simply more complex than other forms of consumer entertainment, and as video game consumers continue to become more sophisticated they will demand more sophistication from video game critics. The solution is for video game critics to draw from the emerging field of game studies. My proposed L/N scoring system is the first step toward applying game studies research to video game review scores.

Labels: , , ,

Saturday, December 15, 2007

Bioshock vs. System Shock

In closing my last post on cognitive games, I proposed comparing the 1994 PC classic System Shock with its spiritual sequel Bioshock, released this year. Both are great games, but are representative of two different philosophies -- and perhaps even eras -- of game design.

To begin with, Kieron Gillen did an excellent job describing just what is so great about Bioshock in a recent article for Eurogamer. I'd encourage everyone to read his article. In fact, I was inspired somewhat by his comparison of Bioshock with System Shock 2. Since I have not played the second System Shock, I will ignore most of the specific points Gillen made in his comparison and draw my own. The main point to be taken from Gillen's article is the excellence of Bioshock's narrative and setting, especially the way in which the game's setting drives the narrative. This is important because placing the narrative within the environment is a useful technique for providing narrative without detracting from the game's interactive nature. As Gillen points out, the more observant and curious the player, the stronger the narrative becomes. In this sense Bioshock's narrative, at least, fits my description of a cognitive game, and Bioshock probably does this better than any game to date.

Where Bioshock falls down compared to System Shock is the way the narrative is tied to the actual gameplay. Both games are littered with recorded messages from their settings' past inhabitants, but while these messages do an excellent job of driving Bioshock's narrative, they have very little impact on the game itself. In System Shock the player needs to listen carefully to these messages in order to figure out what to do. They are clues not only of what happened on the space station and of the people who once inhabited it, but also of what the player must do in order to win the game. Exploration fuels the narrative which then fuels the gameplay itself. Whereas in Bioshock the player only needs to keep pushing forward. With the exception of a few simple fetch quests -- the necessity of which are broadcasted to the player loud and clear -- Bioshock doesn't require exploration and there are no significant puzzles standing between the player and the final credits. As an illustration, I actually played Bioshock with the 'objective arrow' on for most of the game, despite my love of self-guided exploration. Before you cry foul at my apparent hypocrisy, the reason I did this is because I quickly realized that there was no real point to floundering about lost in Rapture when there was always one place you were supposed to be going. The arrow will guide you through almost every part of every level in the proper linear order, and with no nonsequential puzzles in the way or clues to discover, there is no real point in not using it. Either way, you can still take your time and explore the scenery to find hidden item stashes and narrative bits.

The lack of puzzle elements and required exploration alone is not enough to disqualify Bioshock as a cognitively demanding game, however. Bioshock was always billed as a shooter first and foremost and its action elements are more important to it than action elements are to System Shock. Unfortunately, although Bioshock has a great deal of depth in this area, it once again fails to meet my cognitive criteria. A big reason is Bioshock's implementation of vita chambers. Without real consequences in the game, there is less incentive to explore the actual depth that does exist in Bioshock's combat. Furthermore, most of the tactical options that do exist are either too heavy handed to be satisfying(oh look, another huge oil slick to lure enemies onto and set on fire) or impossible to derive deductively(does it make any sense that taking pictures of security cameras would eventually allow you to walk right past them undetected?). To be sure, there are a lot of weapons and skills in Bioshock, but are there many reasons to use one over the other besides running out of ammo?

System Shock had a better implementation of vita chambers because each chamber was inactive until you found it and flipped a switch nearby. This created a mini objective for each level and required you to explore carefully for a good portion of each level. So even though System Shock was not as combat intensive as Bioshock, its combat was still a more challenging experience.

Bioshock is not alone among recent games when it comes to its flaws. In fact, it is quite representative of recent trends in game design, only most recent games don't have such an exquisitely crafted setting to make the whole experience still worthwhile. Games today are targeted at a more casual audience and so developers are focused on creating games with shorter learning curves and less challenging gameplay. These types of games are still entertaining, in much the same way as other forms of pop entertainment, but after 6 to 8 hours the experience becomes stale. From reading forum posts online, I gather a lot of people felt this way about Bioshock. The game just got boring and repetitive to play by sometime around the halfway point. But at the same time, many gamers don't feel that a game is worth the initial $50-60 without 15+ hours of content to play through.

So should fans of cognitive games be without hope? Not entirely. The audience of casual gamers is expanding, but that doesn't mean the cognitive niche is shrinking. The days when cognitive games top the sales charts may be over, but in absolute terms I think they will continue to sell as well as they always have. The differences will be in who is making them and in how much money it costs to make them compared to the sales leaders.

Labels: , ,

Wednesday, December 05, 2007

Cognitive Games

As I've begun to think more about the types of games I most enjoy, I've realized that the games I like best are those that require an investment from the player. It seems that the games which demand the most often have the most to give back. Of course not all demanding games are richly rewarding, but for the most part games that are demanding seem to perform the unique role of games as learning tools better than those which aren't. The enjoyment an individual derives from games compared to other entertainment media depends in part on how much they enjoy the learning aspect of games, and the purer the game, the more it relies on learning.

I call games which stress learning cognitive games. There are two broad categories of cognitive games, complex games which demand the player to learn complicated systems and interfaces, and games with relatively simple mechanics but which require total mastery of those mechanics. In the case of the simple games, the learning process is often more akin to learning to play a sport than it is to the process of learning a complex game. I enjoyed both types of games starting at an early age.

My first exposure to complex games was through pen and paper role playing games like Dungeons and Dragons. P&P RPGs are usually quite complex because their rulesets need to be able to handle an extremely open-ended game and thus adjudicate an endless number of possibilities. Of course, the game master is there to make judgments on how to apply the rules, but the fact that D&D has been so successfully translated to computer simulations which lack human game masters is a testament to the game's complexity. In fact, the computer RPGs which are directly based on P&P systems are probably more complex on average than those which aren't.

In addition to ridiculously complex RPGs, during the 80s and 90s computers also held host to numerous other cognitively demanding games, including turn-based strategy and war games, flight simulators, real-time strategy(RTS) games, and various simulations. Most of these games had thick paper manuals and complex controls which took advantage of every input device the PC had to offer. These are games which required study, and people like me were happy to spend hours poring over the manuals and devising strategy when not actively playing.

Console games, on the other hand, tended toward the simple variety, but there were still cognitively challenging games. Tetris and other similar block puzzle games have very simple systems that are easy to understand but demand faster and faster mental processing and reaction from the player as the difficulty increases. Many side scrolling platformers and action games likewise have fairly simple mechanics(alther much more complex than Tetris) but require dedication from the player to acquire the necessary hand-eye coordination to complete the game. And to provide one more example, fighting games like Street Fighter require the development of timing, coordination, and dynamic strategic thinking. And with each game employing different systems, skill in one doesn't translate 1:1 into other games no matter how superficially similar they may seem.

The one characteristic all cognitive games have in common is that they punish you for making a mistake. Negative feedback is necessary to enforce learning. The trend these days is toward positive feedback in the form of rewards or unlockables, and although those can be useful and fun, I still think cognitive games need negative feedback. If the game is so lenient that the player can progress the narrative even while playing badly, he probably isn't even aware that he is playing badly. Imagine learning to play chess against a computer AI simulating a typical 10 year old. Despite the depth and potential complexity of the game, you would be unlikely to appreciate the finer points without a more challenging opponent. Perhaps one reason competitive multiplayer modes are so popular is because in the modern narrative-focused game that is the one mode where failure, and thus learning, is unavoidable.

I've recently been playing three games which I consider to be good examples of cognitive games, Armored Core 4, Virtua Fighter 5, and System Shock. The telling point to be made here is that System Shock is over ten years old and the other two games are continuations of series which began long ago(Armored Core 4 is actually the 12th game in that series). Are games today less cognitively demanding than the games of old? Are new franchises of cognitive games likely to find mainstream success in today's commercial environment? The recently released Bioshock, a spiritual sequel to the System Shock series, is an excellent subject for analysis, but one which I shall have to tackle in a later post!

Labels: , ,