Friday, July 3, 2009

Another half-year in parses

While others hoard screenshots, I hoard parser files. Another six months, another excuse for a filler post based on parser "output." The point of this exercise is to show that parsing can be a useful summary of your activities and, in some cases, help to assess how well you are doing in aspects of the game other than mindless merit damage.

Edit (July 4): updated

An Affable Adamantking? (June 26)

Damage Summary
Player Total Dmg Damage % Melee Dmg WSkill Dmg Spell Dmg
BLM (me) 5 0.04 % 0 0 5
BLM 1078 8.12 % 0 0 1078
DRK/DNC 2109 15.88 % 1287 526 296
DRK/NIN 10090 75.97 % 9756 0 334
Total 13282 100.00 % 11043 526 1713

Melee Damage
Player Melee Dmg Melee % Hit/Miss M.Acc % M.Low/Hi M.Avg
DRK/DNC 1287 61.02 % 12/2 85.71 % 86/134 107.25
DRK/NIN 9756 96.69 % 83/6 93.26 % 5/187 120.43
Comments: I responded to a Whitegate shout for one of the "beastmen helm" quests that hardly anyone cares about. I had done this previously with NIN/WAR and the assistance of a RDM/WHM by zoning Diamond Quadav until it was isolated from its stooges, and I was interested if they would take a different tack. Actually, their approach called for a DRK-zerg of Diamond Quadav, leaving the BLMs to preoccupy (sleep) the others, which isn't a bad idea yet they still ran out of steam. As you can see, the damage output seemed to be decent enough to pull this off. (The DRK/DNC didn't 2-hour for some reason.) Diamond Quadav being a WHM, of course Benediction ruined this attempt, especially with no attempt to separate the boss from its minions.
Damage Summary
Player Total Dmg Damage % Melee Dmg WSkill Dmg Spell Dmg
BLM (me) 3347 19.23 % 0 0 3347
BLM 5610 32.23 % 0 0 5583
DRK/DNC 906 5.20 % 439 0 467
DRK/NIN 7544 43.34 % 2843 2659 2042
Total 17407 100.00 % 3282 2659 11439

Melee Damage
Player Melee Dmg Melee % Hit/Miss M.Acc % M.Low/Hi M.Avg
DRK/NIN 2843 37.69 % 198/74 72.79 % 0/167 13.44

Weaponskill Damage
Player WSkill Dmg WSkill % Hit/Miss WS.Acc % WS.Low/Hi WS.Avg
DRK/NIN 2659 35.25 % 12/0 100.00 % 38/525 221.58
- Vorpal Blade 2659 100.00 % 12/0 100.00 % 38/525 221.58
With Blood Weapon now unavailable, the only realistic tactic was to isolate Diamond Quadav and proceed to plink away at it with nukes and letting the DRK/NIN "tank." Unfortunately, the guy who wanted to "upgrade" the quadav barbut died without reraise and, in fact, Diamond Quadav is rather accurate for an easily-enfeebled NM, giving the DRK/NIN some trouble with shadows, so I just kited it with gravity and bind until the guy returned along with someone else on bard, making blink-tanking realistic. Meleeing was just terrible (not sure why there was a switch to 1-handed sword), but whatever gets the job done...

Farming Royal Jelly (May 7)

Experience Rates
Number of Fights : 234
Date : 5/7/2009
Party Duration : 14:14:40
Total Fight Time : 2:11:13
Avg Time/Fight : 219.15 seconds
Avg Fight Length : 33.65 seconds

Item Drops
89 beehive chip
13 serving of royal jelly
22 insect wing
6 giant stinger
Comments: For those aspiring to level cooking to 100, it's either Red Curry or Cursed Soup, the latter requiring Royal Jelly, which was inexplicably flagged "exclusive" by some asshole on the "dev team." With a glut of 20 red curries languishing on a mule and sitting somewhere above 99 skill, I tried my hand at farming this shit.

Where do you farm Royal Jelly? You can risk dying to Final Sting while farming pephredos in Wajaom Woodlands (if you melee) or mow down all the Death Jackets, all on a 14-minute respawn timer in Crawler's Nest. 234 bees later, I got a 13th Royal Jelly and I still didn't get to 100 cooking.

More Aura Statues x58 (Jan 21)

Debuff     # Times   # Successful   # No Effect   % Successful
Aspir 2 2 0 100.00 %
Bind 60 43 0 71.67 %
Gravity 117 106 2 90.60 %
Sleep 5 4 0 80.00 %
Sleep II 17 17 0 100.00 %
Stun 37 37 0 100.00 %
Comments: Aura Statues are bothersome with relatively poor enfeebling skill as I showed last time. But at this point, I am pretty sure I had all the key enfeebling pieces, including Oracle's Gloves and Enfeebling Torque, but no Witch Sash, Enfeebling Earring, or corresponding elemental grip. Even so, Gravity outright resisted 9 of 115 times. I would imagine scholar with dark arts and the appropriate equipment and merits would have little trouble enfeebling statues.

Damage mitigation on WAR/SAM with Greater Colibri

Damage Taken Summary
Player Total Dmg Damage % Melee Dmg Abil. Dmg
WAR/SAM (me) 7727 50.27 % 5660 2067
WAR/NIN 3687 23.99 % 1885 1802
DRG/SAM 1944 12.65 % 1304 640
RDM/WHM 145 0.94 % 145 0
BRD/NIN 1432 9.32 % 1432 0
BRD/WHM 437 2.84 % 437 0
Total 15372 100.00 % 10863 4509

Passive Defenses
Player Evasion Evasion % Parry Parry % Counter Counter %
WAR/SAM 4 3.13 % 7 5.65 % 17 15.18 %
WAR/NIN 3 3.41 % 7 8.24 % 0 0.00 %

Active Defenses
Player Shadow Shadow % Anticipate Anticipate %
WAR/SAM 0 0.00 % 62 52.99 %
WAR/NIN 56 71.79 % 0 0.00 %
DRG/SAM 0 0.00 % 3 23.08 %
BRD/NIN 54 85.71 % 0 0.00 %
Comments: Where Utsusemi is involved, in KParser the "blink" rate is the number of absorbed attacks over the total number of attacks that weren't evaded or parried. The total number includes TP moves (but not their individual hits). From experience with Greater Colibri, my rate has ranged from 82% to 90%, which seems acceptable if not a sign of hyper-vigilance in recasting Utsusemi.

Similarly, a "Seigan rate" can be calculated by substituting the sum of counters and anticipates for the number of blinked attacks. To the extent that the Seigan rate, as a measure of "active" defensive efficiency, can be maximized with judicious use of Third Eye, it could help identify room for improvement. I have yet to see any discussion about what an optimal Seigan rate might be, though.

Unfortunately, my personal insight on WAR/SAM damage mitigation consists solely of two pickup parties. The parser output above is a partial record of a short (approximately 30 minutes) March 13 merit party where I was actually allowed to sub /SAM. My Seigan rate was 79/128 = .617.
Damage Taken Summary
Player Total Dmg Damage % Melee Dmg Range Dmg Abil. Dmg
WAR/SAM (me) 16069 45.15 % 13888 0 2181
BRD/NIN 4669 13.12 % 4669 0 0
DRG/NIN 2780 7.81 % 1576 0 1204
BRD/WHM 176 0.49 % 176 0 0
WAR/SAM 11527 32.39 % 9227 0 2300
WHM/SCH 372 1.05 % 372 0 0
Total 35593 100.00 % 29908 0 5685

Passive Defenses
Player Evasion Evasion % Parry Parry % Counter Counter %
WAR/SAM (me) 10 3.02 % 10 3.12 % 40 13.11 %
BRD/NIN 9 4.39 % 0 0.00 % 0 0.00 %
DRG/NIN 4 2.37 % 10 6.06 % 0 0.00 %
WAR/SAM 9 4.84 % 9 5.08 % 16 9.94 %

Active Defenses
Player Shadow Shadow % Anticipate Anticipate %
WAR/SAM (me) 0 0.00 % 166 53.38 %
BRD/NIN 164 83.67 % 0 0.00 %
DRG/NIN 134 86.45 % 0 0.00 %
BRD/WHM 3 60.00 % 0 0.00 %
WAR/SAM 0 0.00 % 85 50.60 %
This parser output summarizes defensive efficiency from a June 17 polearm-only (hey, I was curious) merit party (82 minutes) where I was also allowed to sub /SAM. My Seigan rate was 206/311 = .662, a sign of personal improvement but also an indication that I could be more efficient as I was being rather lazy with Seigan renewal. The other WAR/SAM had a Seigan rate of 101/168 = .601. In contrast, the DRG/NIN actually outparsed both of us (draw your own conclusions) and was much more efficient defensively.

Innin could've come in handy here (April 16)

Damage Summary
Player Total Dmg Damage % Melee Dmg WSkill Dmg
NIN/WAR 4454 100.00 % 2854 1600

Melee Damage
Player Melee Dmg Melee % Hit/Miss M.Acc % M.Low/Hi M.Avg
NIN/WAR 2854 64.08 % 87/62 58.39 % 11/43 30.64
Comments: The reaction to the upcoming (July 2009) Ninja job adjustments, particularly Innin, a new ninja ability that "lowers enmity in exchange for reduced evasion" while conferring bonuses to accuracy, critical hit rate, and ninjutsu damage "when striking your target from behind," has been mixed to put it charitably.

This combination of lower enmity (sounds like this will have the effect of lower rate of enmity increase with Innin active) and increased damage-dealing capability seems peculiar, especially since these new abilities are subject to decay per "development" team fetish. I can imagine the enmity change doesn't "decay" while the damage-dealing part does. But it could be of use in low-number activities where ninjas can still melee, provide enfeebling support, and don't have to worry about positioning, but basically concede "tanking" to a far superior DD (like monk), possibly in conjunction with thieves being unaccountable for the damage they inflict. (Oh, you call that hate control?)

The above output from an "arena-style" fight from the quest "Bonds That Never Die," is a partial picture of how feeble (my) ninja was even with sushi. Being totally outclassed by 2-handers (samurai) who ended up tanking the latter half of the fight, Innin could've helped to speed up the fight.

Okay, this is a really weak argument for Innin, but at least it's something.

An example of NW Apollyon soloing failure (May 10)

Fight #   Enemy               Start Time   End Time   Fight Length
1 Bardha 11:55 AM 11:57 AM 00:02:27
3 Mountain Buffalo 12:00 PM 12:05 PM 00:05:03
4 Mountain Buffalo 12:05 PM 12:13 PM 00:07:32
5 Apollyon Scavenger 12:15 PM 12:17 PM 00:01:40
7 Gorynich 12:20 PM 12:22 PM 00:02:14
8 Gorynich 12:22 PM 12:24 PM 00:01:32
9 Gorynich 12:26 PM 12:28 PM 00:01:42
10 Gorynich 12:31 PM 12:33 PM 00:01:54
11 Gorynich 12:36 PM 12:37 PM 00:01:38
12 Kronprinz Behemoth 12:43 PM 12:47 PM 00:04:16
13 Kronprinz Behemoth 12:50 PM 12:52 PM 00:02:26
14 Kronprinz Behemoth 12:55 PM 12:58 PM 00:02:46
15 Kaiser Behemoth 1:01 PM 1:32 PM 00:31:11

Player Spell Dmg Spell % #Spells S.Low/Hi S.Avg
- Aero IV 629 7.47 % 1 629/629 629.00
- Bio 5 0.06 % 1 5/5 5.00
- Bio II 434 5.15 % 8 16/69 54.25
- Blizzard IV 6345 75.33 % 8 770/853 793.13
- Drain 1010 11.99 % 9 35/165 112.22
Comments: This output represents a wasted opportunity to clear NW Apollyon with ease on my first legitimate attempt. You can see where I wasted a lot of time even with only two buffaloes to kill. Also, I lost valuable time by dying to Kaiser Behemoth somehow. Aspir accuracy is indeed a rate-limiting step (NW Apollyon motivated my earlier posts on Aspir), so to speak, and for some reason I managed to cast Aspir only three times in a half-hour (47, 91, 81). It seems Kaiser Behemoth can be killed in 20 minutes solo (unless the video from some jackoff hume BLM I saw was subtly sped up), so this was very disappointing to me. It's one thing to know what to do, and another to execute actually.

In later attempts, I also decided to cast three times in a "lap" around the fifth floor (Blizzard IV, Aspir, and Bio II or Drain where applicable) as opposed to the two times that others do, but this seemed counterproductive since I wasn't really inflicting damage at a faster rate with a "three-point" approach and I was exposing myself to higher risk.

Ninja, Marinara Pizza, and Temenos - Western Tower

Damage Summary
Player Total Dmg Damage % Melee Dmg WSkill Dmg Spell Dmg Other Dmg
PLD/NIN 32083 14.21 % 25625 5980 0 371
NIN/WAR (me) 79879 35.37 % 53089 26707 0 83
DRK/NIN 53584 23.73 % 28735 21994 1960 895
RDM/WHM 7327 3.24 % 0 0 7327 0
THF/NIN 51155 22.65 % 30764 19718 0 134
SC: Detonation 833 0.37 % 0 0 0 0
SC: Scission 965 0.43 % 0 0 0 0
Total 225826 100.00 % 138213 74399 9287 1483

Melee Damage
Player Melee Dmg Melee % Hit/Miss M.Acc % M.Low/Hi M.Avg #Crit C.Low/Hi C.Avg Crit%
PLD/NIN 25625 79.87 % 672/121 84.74 % 0/77 35.50 56 32/112 67.11 8.33 %
NIN/WAR (me) 53089 66.46 % 1069/145 88.06 % 0/90 43.69 181 32/139 78.98 16.93 %
DRK/NIN 28735 53.63 % 252/57 81.55 % 0/222 104.41 29 138/287 187.97 11.51 %
THF/NIN 30764 60.14 % 797/72 91.71 % 0/68 28.57 140 29/381 85.67 17.57 %

Weaponskill Damage
Player WSkill Dmg WSkill % Hit/Miss WS.Acc % WS.Low/Hi WS.Avg
NIN/WAR (me) 26707 33.43 % 46/0 100.00 % 159/901 580.59
- Blade: Jin 26078 97.64 % 44/0 100.00 % 268/901 592.68
- Blade: Kamu 629 2.36 % 2/0 100.00 % 159/470 314.50

Passive Defenses
Player Evasion Evasion % Parry Parry %
BRD/WHM 1 11.11 % 0 0.00 %
PLD/NIN 13 8.28 % 3 2.08 %
NIN/WAR (me) 25 11.26 % 6 3.05 %
THF/NIN 3 15.00 % 0 0.00 %

Active Defenses
Player Shadow Shadow %
PLD/NIN 76 53.90 %
NIN/WAR (me) 152 79.58 %
DRK/NIN 3 27.27 %
THF/NIN 5 29.41 %
Comments: I had a chance to do Temenos West as NIN/WAR, so I took this opportunity to see how well I could do with Marinara Pizza (+1), which is kind of a boon for 1-handed melee since it overcomes the accuracy deficit that 1-handers face compared to 2-handers and also provides an attack bonus.

Even 44+ accuracy (estimated) wasn't enough to achieve maximum hit rate, though.
Damage Summary
Player Total Dmg Damage % Melee Dmg WSkill Dmg
RDM/WHM 4133 1.07 % 0 0
WHM/SCH 1395 0.36 % 0 0
THF/NIN 62995 16.29 % 35352 26928
NIN/WAR 110655 28.61 % 78434 32156
WAR/NIN (me) 113930 29.45 % 72330 41287
PLD/NIN 49607 12.82 % 29385 19984
Diabolos 231 0.06 % 231 0
Garuda 34906 9.02 % 5926 0
Leviathan 72 0.02 % 72 0
Shiva 1615 0.42 % 166 0
SC: Detonation 2272 0.59 % 0 0
SC: Impaction 307 0.08 % 0 0
SC: Light 2034 0.53 % 0 0
SC: Reverberation 1676 0.43 % 0 0
SC: Scission 978 0.25 % 0 0
Total 386806 100.00 % 221896 120355

Melee Damage
Player Melee Dmg Melee % Hit/Miss M.Acc % M.Low/Hi M.Avg #Crit C.Low/Hi C.Avg Crit%
THF/NIN 35352 56.12 % 897/104 89.61 % 10/65 30.99 178 28/290 73.42 19.84 %
NIN/WAR 78434 70.88 % 1510/99 93.85 % 0/103 45.94 231 27/150 85.17 15.30 %
WAR/NIN (me) 72330 63.49 % 458/25 94.82 % 50/247 143.30 67 135/306 243.28 14.63 %
PLD/NIN 29385 59.24 % 797/204 79.62 % 0/88 35.23 32 42/128 76.13 4.02 %

Weaponskill Damage
Player WSkill Dmg WSkill % Hit/Miss WS.Acc % WS.Low/Hi WS.Avg
NIN/WAR 32156 29.06 % 55/0 100.00 % 220/1023 584.65
- Blade: Jin 30307 94.25 % 50/0 100.00 % 220/1023 606.14
- Blade: Kamu 1849 5.75 % 5/0 100.00 % 259/490 369.80
WAR/NIN (me) 41287 36.24 % 56/0 100.00 % 161/1615 737.27
- King's Justice 35740 86.56 % 50/0 100.00 % 161/1291 714.80

Passive Defenses
Player Evasion Evasion % Parry Parry %
RDM/WHM 1 7.14 % 0 0.00 %
THF/NIN 4 50.00 % 0 0.00 %
NIN/WAR 42 17.00 % 4 1.95 %
WAR/NIN (me) 4 5.13 % 2 2.70 %
PLD/NIN 16 11.76 % 5 4.17 %

Active Defenses
Player Shadow Shadow %
THF/NIN 3 75.00 %
NIN/WAR 165 82.09 %
WAR/NIN (me) 56 77.78 %
PLD/NIN 77 66.96 %
This other (successful) attempt to clear Temenos West involved a different NIN/WAR using Dorado Sushi. It was really annoying to see average melee and WS damage similar to mine when using Marinara +1, but this could be explained by other factors (birds have low defense, Usukane, more katana merits, more consistent application of Dia II, etc.).

Friday, June 26, 2009

Retaliation

The warrior job ability Retaliation (level 60) is something I rarely use outside Nyzul (for bosses), so I was curious about how frequently it activates. This is something I would probably have a good idea about just from doing Campaign... if I actually did Campaign. Of course, FFXIclopedia's article is not much help as to what actually affects this frequency (there seems to be variability based on something else other than neglecting to reactivate Retaliation), so I went looking for any sort of data elsewhere.

From wiki.ffo.jp, apparently it is thought that Retaliation rate is based on weapon delay, from 20% at 999 delay to 50% at 200 delay, subject to further verification. (At least this is what I gleaned from Google Translate.)

As for data to support that kind of trend, so far I found one source that describes the results of dual wielding Brass Jadagna +1 (334 delay) and Caduceus (216 delay). Indeed, it is not immediately obvious whether Retaliation, in a dual wield context, is dependent only on the main weapon or based on total delay. The results indicate that Retaliation rate depends only on the main weapon:

Brass Jadagna +1 (main)/
Cauduceus (sub): 71/244 (.291) with 95% confidence interval (0.2347960, 0.3523304)

Caduceus (main)/Brass Jadagna +1 (sub): 110/241 (.456) with 95% confidence interval (0.3923515, 0.5215964)

I would assume the experiment was done with Retaliation constantly active, at maximum hit rate (95%), and the Retaliation rate is not dependent on mob's melee attack speed and level (to account for the possibility that there were different mobs involved). Even more important, I would assume the rates were estimated by taking the ratio of Retaliation procs over the number of times being hit.

If these assumptions actually hold, then Retaliation frequency does appear to depend on the main weapon delay only, not the total delay from both weapons. The higher delay weapon, Brass Jadagna +1, had a significantly lower rate than the lower delay weapon, Caduceus. It is not clear whether dual wielding would be any different from using a 2-handed weapon, but I don't see why there would be a difference.

Chocobo update: winning streak ended at 10; have won 12 of the last 13 (7 uncontested). Record is 111-79-24-15.

Wednesday, June 24, 2009

Phantom Roll and support for discretion to roll however you want

Before I discuss the optimality of several approaches to using Phantom Roll, I want to talk glibly about whether Phantom Roll even involves the use of a fair die in practice, which is the major assumption underlying Phantom Roll "strategies." I will use Pearson's chi-square test to check badness of fit of the following count data.

This is the only discussion I've seen so far (2006) that entertains the possibility that the outcome of Phantom Roll (I through VI inclusive) is not uniformly distributed. The point of the data collection was to find some evidence that the die becomes weighted in the presence of the optimal job associated with the specific roll (Bard with Choral Roll, etc.). You can read the thread for details.

The first example, with 700 uses of Corsair's Roll, did not yield compelling evidence against unbiasedness (p-value .0811).

The other eight examples involved sampling 100 times under varying conditions. At this point it bears reminding that the distribution of p-values under the null hypothesis of a fair die is (asymptotically) uniformly distributed (keeping in mind bin specification for the sake of generating histograms), as illustrated below with a bunch of simulated p-values sorted into histogram bins, given a sample size of 100.


I bring this up only as a reminder of what the possible p-values can be under the null hypothesis.

For the remaining eight data sets, tests for unbiasedness yield p-values of .4614, .2739, .3601, .007439, .7974, .3101, .09696, and .2763. Based on this crude analysis, only the data set for Healer's Roll with WHM present showed statistically significant evidence of biasedness (specifically 30/100 for a roll of I), but compared to the other non-significant results, it seems difficult to attribute this to something other than Type I error.

Of course, the primary question of interest was not whether Phantom Roll gives unbiased rolls regardless of situation, but whether the presence of the optimal job changes the "weight" of the roll. Tests for homogeneity for each specific roll (multiple testing duly noted) give p-values of .5402 (bard), .1077 (white mage), .6433 (ranger), and .1099 (thief).

Generally speaking, chi-square tests have pretty low power, and one tends not to "invert" these to (sets of) confidence intervals to get a good sense of how (in)adequate the sample sizes are. But considering this data as a whole there isn't a particularly good reason to think that the Phantom Roll die is biased.

Now, optimality of two Phantom Roll approaches

The following could basically be summarized as comparing the pros and cons of busting more versus busting less depending on how you go about doubling up.

There is a spreadsheet that provides a convenient summary of whether to Double-Up for various roll types, based on a criterion of conditional expectation (actual buff value), given the current roll total. Basically, consideration of (conditional) expected value is a formal way to make a decision that can be mostly carried out using common sense--you will never double up with a total of 11, as the expected value of the buff after Double-Up must be 0--but addressing borderline cases where it may not be obvious whether one should double-up, for example if your current roll is 6. I awkwardly call this the "expected value on double-up" (EVDU) approach.

The spreadsheet also gives an unconditional expected value of the roll after doubling up based on the expected value criterion, which could be useful for comparing different types of rolls on a "long-run" basis.

For wannabe nerds who can't even calculate the conditional expectations or understand probability, that is one way to go about it. Not unexpectedly, these min/maxing wannabe nerds frown upon conservative approaches that seek to minimize the probability of a Bust, with the implication that people who refuse to Double-Up on a 6 are "suboptimal." For the remainder of this post, I will call categorical refusal to Double-Up on 6 (unless 6 is unlucky), yet still Doubling-Up if one gets an "unlucky" total (therefore risking a Bust), as the "conservative" approach (and the only one I will consider in this post).

I am willing to bet that most of the people who advocate EVDU (implicitly or not) never actually bothered to compare EVDU with more conservative approaches quantitatively, especially for specific types of rolls. By quantitatively, I mean comparing (unconditional) expected values under each approach to see how much better in the long-run EVDU is, and also comparing the busting proportions under each approach to see how much riskier in the long-run EVDU is.

Personally, I don't really give a shit what rolling strategy a corsair actually uses, since to me it mostly falls under the purview of individual playing style.

Consider Corsair's Roll, for example. Under EVDU, the expected percentage increase in EXP is 15.66% while the conservative approach gives an expected increase of 15.55%, which to me is a really trivial difference. Moreover, the probability of busting under EVDU is .051 while the probability of busting "conservatively" is 0. If you are willing to assume an actual 5% (non-zero) chance of busting for a theoretical 0.11% long-run increase in EXP, fine. But here, the tradeoff between risk and reward is not all that good.

I also estimated the probabilities for the Corsair's Roll bonuses under each strategy (since I didn't want to waste even more time thinking about how to hand-calculate them) to make it easier to compare the strategies in probabilistic terms. (Relative frequencies may not add up to 1 due to rounding.)

COR Roll
EVDUConservative
Bust
.051
.000
8%
.134
.082
13%
.000
.309
15%.193.142
16%
.165
.114
17%
.095.044
20%
.309.309
24%
.052
.000

I colored the relevant probabilities one "side" might use to make a case against the other. Note that the probability of obtaining the "lucky" result (20% EXP increase) is the same regardless of approach. I also did the same for Hunter's Roll (melee and ranged accuracy) and Chaos Roll (melee and ranged attack) without the presence of the optimal job.

Again, the tradeoff between risk and reward is not so great. Whether you, as a corsair, want to make that tradeoff should be up to you and not to dumbasses who need to rely on mindless rules of thumb because they don't know any better. Personally, I would rather allocate all of my busting risk to another roll rather than to Corsair's Roll if the increased risk is actually worth it on another roll. But when is it worth it? I repeat the above exercise with both Hunter's Roll and Chaos Roll, rolls that are available early on.

For Hunter's Roll, the expected value under EVDU is 29.63 accuracy, and 28.09 taking the more conservative tack. Clearly, a 1.54-point difference in average accuracy is such a profound increase as to assume a greater risk of busting. The estimated probabilities are given below.

RNG Roll
EVDUConservative
Bust
.135
.057
20
.000
.3o9
25
.194
.142
27.161.101
30
.124
.063
40
.264 .264
50
.122
.064

For Chaos Roll, the expected value under EVDU is 18.6% attack increase (47.5/256), and 17.6% attack increase (45.0/256) playing it conservatively. Again, a 0.98% average difference in attack obviously warrants the increased risk of busting. The estimated probabilities are given below.

DRK Roll
(xx/256)
EVDUConservative
Bust
.134
.058
32
.000
.3o8
40
.193
.142
44
.163.101
48
.124
.063
64
.265 .265
80
.124 .063

Sure, a 1-point or 1% difference may be important enough to you, but 0.11%?

I spent time constructing this post while considering whether to level corsair to 75. (I won't but not based on what I found in this post. Ultimately I'd rather buy an account with a ready-to-play COR75 than waste time leveling another job to 75.) Take-home message: do whatever the hell you want as long as you can support it logically.

Thursday, June 18, 2009

Milestones

Somehow I managed to update sporadically this testament to a lack of priorities for almost one year. To be honest, I wrote foremost for myself so I didn't really bother to spend much time making these entries easily digestible for a wider audience. In particular, I found an excuse to apply some basic probability and statistics to FFXI, the playing of which is also a testament to a lack of priorities and extremely poor taste. On the other hand, I did try to focus my attention on the mechanics of the game so that these entries would have some informative value--at least some people thought so--instead of being just some masturbatory self-chronicle.

I would really like to maintain this conceit, anyway, but there is not much of an empirical mentality among the so-called playerbase to provide persuasive support for any theories that are developed or serendipitously discover non-trivial things about how the game works. Maybe some dead-enders take great pride in running their mouths without putting their bullshit to the test, but I find legitimizing claims with real data and observations to be far more interesting. I thought about turning this blog into a "digest" of sorts to summarize both new and old findings and give credit to the individuals that shed some light on some aspect of game mechanics, but actually it is just too tiresome to comb forums with crap search functions and garbage "intellects" for shiny nuggets of insight. So I will just continue talking about things that interest me enough to commit to blog, even though the posting frequency based on that criterion will be very low.

Anyway, this wasn't supposed to be just some navel-gazing exercise. Instead of making individual posts for the following topics, none of which really warrants standalone status, I decided to throw them all into a single entry.

Thoughts on Aspir

I managed to finish collecting some data to check the effect of Pluto's Staff and magic accuracy +12 on the potency of Aspir and updated the dot plot:

I think it's safe to say that INT or magic accuracy (MAB too, based on tarblm's results) have no role in the potency of Aspir (and by analogy, Drain). Note that I am not bothering with statistics and just arguing informally that none of MAB, macc, and INT increased the maximum in these samples.

As far as accuracy is concerned, that is much more inconvenient to check. The main purpose of my collecting data was to visualize the distribution of Aspir values. It seems here that the range of possible values of unresisted Aspir is fairly wide. The low values of Aspir observed may also be the result of a half-resist, which may cut the unresisted Aspir value in half. Unfortunately, you can see how partially resisted Aspirs are easily confounded with unresisted Aspirs if these ideas are true. Going back to tarblm's data though, the observed Aspirs generally have much greater variability, which could be attributed to partial resists.

Chocobo racing

Haven't talked about this in a while. A few months ago I actually canceled my content ID, but I let myself get pulled back into this pit of mediocrity that is FFXI. Since then I made some effort to maintain more detailed information on my Chocobo Circuit results, particularly whether my chocobo was competing against other PCs.

Not surprisingly, few of my C1 races were uncontested. In fact, 16 of the first 20 after I returned had at least one PC chocobo and 7 of those 16 had 2 PCs. I won only nine of those races, with a pretty abysmal 3-3-1-1 record against only one other PC.

During this time I came across a testimonial of another chocobo racer, also with a SS/B/B/B chocobo, who claimed to have won 67% of his races (128-41-22). This kind of pissed me off because farming chocobucks is extremely boring and here this guy was getting nearly 2 million more gil with the same chocobo profile and similar number of races. I thought perhaps he faced much less competition on his server and that his use of leather saddles may also have been a factor in his great success. But rather than shell out another $25 bucks to SE just for a server transfer, I tried the Sheep Leather Saddle for another month to see if taking the receptivity hit would be worth it in races with one or more other PCs.

My results were even worse with the leather saddle. In 20 races, I went 8-10-0-2 with 11 contested races, and in four of those contested races, an NPC chocobo won (I placed 2nd in all four) and in one case, two NPCs actually placed 1-2. (I finished out of the top 3 in this one.) Even worse, I won only four of the uncontested races, races in which I really "needed" to win to blunt the annoyance of losing chocobucks.

I then went back to the elm saddle and am now nurturing a nine-race winning streak (four contested), by far the longest streak I ever had. During this streak, I also reached the 10-million gil mark in net earnings. I am now going out of my way to race only in "off-hours" time slots to try to get my win rate back to 50%. My current record is 107-78-24-15.

Having reached 10 million in net earnings, I calculated an approximate rate of gil per hour earned in chocobo racing based on the time (one free race per 5 minutes) and gil spent to accumulate enough chocobucks (5,846) to enter the races and the gross earnings. Including this nine-race winning streak, chocobo racing has yielded an average of 74,644 gil per hour. (I admit this figure does not include time spent running to Chocobo Circuit.) Not as efficient as the guy winning 67% of his races, but still a nice reminder that even though chocobuck farming is a real pain in the ass, at least this huge barrier to entry allows chocobo racing to provide a steady income to those who actually put up with it... assuming the C1 races aren't so congested.

Thoughts on the possibility of chain 6 solo on Ebony Puddings without Novio and Manafont

I was extremely disappointed to find out that I had died 957 times between Adventurer "Appreciation" 2008 and A.A. 2009, the majority on black mage. (If I am really appreciated, can I get a Chocobo Wand in fewer than two weeks without being lucky?) I had considered myself more risk-averse over the past year (I did not even do Dynamis at all), but in retrospect this was not true, since you tend to die a lot in pickups, soloing, and poor event linkshells.

The pain of losing experience points on black mage (never mind the bullshit conceit that losing experience points on top of losing the time and resources you wasted is a reasonable penalty) can be blunted somewhat with efficient rate of gain of EXP. But what is considered efficient? From experience, the best I can do is around 8,000 EXP/hour (estimated by the time required to burn off an Emperor Band), and that's when being somewhat vigilant about achieving chain 5 (which is trivial if you are paying attention). Anyone who claims rates of 10,000 EXP/hr solo is full of shit until proven otherwise.

Is it possible to achieve chain 6 solo without Novio, though? By the time chain 5 rolls around I almost always have insufficient MP without Aspir to mount a chain 6 attempt, an indication that my maximum MP is not high enough. Moreover, even three "tier 4" nukes tend to leave a sliver of HP (on off-weather days), meaning that I have to rely on Drain to finish off a pudding. Casting Drain on chain 4 and 5 wastes MP that could be used for chain 6. These factors, along with half-resists and weather effects, conspire to make it really difficult to achieve chain 6. If it is easier than I am pondering, I'd like to know though, but not from shit Morrigan's users who still cast AM II on puddings.

Benevolent Despot

With the advent of Fields of Valor, the prospect of not spawning Despot in a timely fashion is less unpalatable since a training regime awards 1,550 EXP in about an hour of killing 11 placeholders. And maybe those tabs will actually be useful someday. Also, soloing without desirable rewards has grown pretty tiresome. I still have a Fenrir solo (ninja) on the back burner now that Lunar Roar supposedly does not dispel reraise, but I haven't been motivated to do that yet. At the moment, Despot is the only remotely appealing "get out there and kill shit" profit opportunity with a relatively high barrier to entry (actually being able to solo it under 30 minutes to lower the chance of vultures finding you and trying vainly to MPK you).

Though not quite as enjoyable as killing Despot and hardly a consolation prize, watching hapless groups kill Despot can provide some humor to brighten the day. Not a few weeks ago, I had the "pleasure" of witnessing this quartet of PLD/NIN, NIN/DRK, BLM and MNK struggle with Despot for over 40 minutes! Apparently, it didn't occur to these people to shed enmity via teleport in order to expedite the kill. Unfortunately, even inept players win eventually, a testimony to the dominance of the lowest common denominator in FFXI. Throw infinite time and resources at something and you can triumph! (Except for Absolute Virtue.)

Spending time farming gems of the west also has given me an opportunity to "get back" at the MPKing Tarutaru Duo That Shall Not Be Named by dispatching Despot in 22-27 minutes while those oblivious assholes continue to kill placeholders long after I'm gone.

Thursday, June 4, 2009

Aspir data and observations

Edit (June 5): some attempts at clarification.

What affects the accuracy of dark magic? Skill only? How about potency? How can we describe the distribution of MP absorbed with Aspir? What data are out there to support the prevailing assertions? Finding old Aspir data sets (from 2006) was easy enough, and I also collected some Aspir data on my own.

This may seem like treading old ground if not for the B.S. I cited earlier in the week. At least there is some data you can cite when making an argument now.

Regarding the 2006 data set, the writer (whom I will call "tarblm" for now) collected Aspir data from King Buffalo (lv 79-82) over six trials. Each trial involved only a single buffalo. The data were collected under one of three configurations, all with a Pluto's Staff:
  • "Dark magic skill": +40 dark magic skill (above 269) was the primary factor
  • "MAB": +30 MAB from equipment (relative to control) was the primary factor
  • Control: 269 dark magic skill
The writer also noted the experience gain for each buffalo. The data are presented in a dotplot:


First, the data give the impression that some amount of dark magic skill increases the average MP absorbed, whereas MAB doesn't, confirming previous beliefs. Certainly the maximum Aspirs are higher. Low values of MP drained seem fairly rare and set apart from the rest of the data, so describing the data as coming from a uniform distribution doesn't quite work.

Was "effective" magic accuracy capped on "very tough" buffalo? I would say yes. Otherwise, level difference would have confounded the results. If there is no difference in accuracy across the trials, one could attribute the average to an increase in so-called potency alone.

To try to avoid that uncertainty about capped magic accuracy for my data, I focused my attention on low-level Tunnel Worms and collected 50 Aspir samples for each of the following conditions without a Pluto's Staff:
  • Control: 77 INT, 269 dark magic skill
  • INT: +43 INT above control (120 total)
  • Dark magic skill: +22 dark magic skill above control (291 total)
Since the worms are so low in level, I just assumed my effective magic accuracy was capped. This is a major assumption but a reasonable one given magic skill level. (Correction: I earlier used a level correction argument to make this assumption. It has not been established that level difference affects mobs in the same way that it affects PCs.) The data are illustrated in the following dotplot:


Similar to tarblm's data, some amount of dark magic skill seems to increase the average MP absorbed, although the increase is not statistically significant. As magic accuracy was probably capped in this scenario, it is probably safe to say that dark magic skill would increase potency by a statistically significant amount if I had more dark magic skill to pile on. In terms of the distribution of Aspir, it seems to shift the range of possible values to the right.

INT doesn't seem to cause any change in potency. That is not to say INT doesn't affect accuracy in some way! Low values of MP (here, below 50) were infrequent and I assure you they didn't result from capping total MP. A decrease in accuracy may manifest in a higher frequency of low values, resulting in a lower average if not a shift in the range of possible values.

Note that last time I gave an example of a data set (on FFXIclopedia) that showed INT increased the average Aspir, and I said this was a potency-only effect based on the assumption of capped accuracy. Perhaps there are other confounding factors that were not cited.

Also, data overall give an impression that Pluto's Staff affects potency (one, the 2006 data are more variable, Aspirs achieve higher values despite King Buffalo being 62-72 levels higher than Tunnel Worm).

From all this, it seems reasonable to conclude that
  • Dark magic affects potency (still not sure about magic accuracy attribute, e.g., from equipment)
  • MAB does not affect potency
  • INT does not affect potency but it seems likely to affect accuracy in some way. Do not confuse accuracy with potency. In a sense, increasing one or the other should still increase the average drained up to a point, but the way each does that is different. An analogy to melee attack and melee accuracy should make sense.
  • Magic accuracy probably affects accuracy but not potency, but I haven't found nor collected any data to check this.
You may also have noticed that the maxima and sample means in the 2006 data set are larger than those in the 2009 set. I'm pretty sure the difference can be attributed to Pluto's Staff.

Tuesday, June 2, 2009

Dark magic and INT

This is just going to be a quick and dirty post but may motivate a more enlightened post later, but just to reinforce the ignorance exhibited by the "player base," here's another chuckle-inducing, basically worthless discussion about what affects the accuracy and potency of Aspir. Here's another discussion from idiots talking about "tests" on the potency of Aspir, but where's the fucking data? Here's an obvious question. Did anyone ever actually disentangle accuracy from potency in a controlled experiment? Or how about, if you are at some hypothesized potency cap, why would you expect potency to increase with anything?

I mean, really, assertions like "Accuracy of [Aspir] is most highly affected by Dark Magic Skill, and is not affected by Magic Attack Bonus, or INT" are based on something rather than pointless anecdote, right?

Actually, in the talk discussion of that FFXIclopedia article on Aspir, there seems to be something like a controlled experiment with random sampling, with a RDM75 (200 dark magic skill) casting Aspir on a single worm between level 10 and 12. Assuming that "accuracy" is capped in some sense on such a low-level target, it appears that INT does affect potency (just do a quick two-sample t). Sure, type I error, blah blah, but it's better than total bullshit.

Tuesday, May 26, 2009

pDIF and obsession with polynomial fits

A recent thread on the BG forums about an investigation of "cRatio for two handers" really underscores the ignorance coming out of the "playerbase" that actively chooses to post on forums. To wit:
  • You got some motherfucker implying the "center" of advanced knowledge about game mechanics lay among those who were banned for Salvage duping, when the "center" is actually maybe 10 people at best, and not all are necessarily English-speaking, much less posters on BG.
  • Someone rightly points out that said motherfucker is some obsequious cock-gobbler (ok, those are my terms) since information on pDIF has been outdated since the "2-hander update" (well before the bannings) yet no one has actually bothered to do an honest investigation.
  • Another one actually bothers to collect some data on damage frequency to see what kind of distribution the data follows, but is easily derailed with a fetish for polynomial fitting to the data and data following a normal distribution (polynomial fitting and normality are contradictions as I will discuss soon).
This inexplicable obsession with polynomial fitting and normality is misguided for several reasons:
  • Normal distributions have obvious tails at the extremes. Moreover, the tails are neither too short nor too long. The data do not show evidence of any real tails.
  • Normal distributions are not parameterized by extrema (minimum and maximum). The parameters are the mean (center) and variance (spread).
  • A second-order polynomial fit cannot "account" for tails. This is obvious because normal distributions have inflection points. So you cannot use a polynomial fit and argue for "normality" at the same time.
  • Coefficient of determination can be thought of as a summary of a model fit. It doesn't mean the model is actually good. You can draw a squiggly line through all the data points and that will give you a R2 of 1, but that would be a terrible and useless model. Polynomial fitting is similarly terrible and useless for the above reasons.
  • Why even bother with any kind of fitting? As long as the distribution is symmetrical, at least you know the minimum, maximum, and median (same as mean) for pDIF given some value of cRatio, so you can use an expected value argument for long-run damage.
That said, there were some useful comments about the "shape" of the data. One poster actually suggested the data may follow some trapezoidal distribution. This is actually quite plausible under probability theory!

Obviously, the data do not appear to follow a uniform distribution. Even acknowledging the discreteness of damage (due to rounding) such that the minimum and maximum might be observed rarely, a uniform distribution of pDIF (NOT DAMAGE) is not all that likely. Although we cannot observe the uniform distribution of pDIF directly, we can observe histogram of damage (NOT pDIF). For this histogram, if pDIF were actually uniform, one might see an extreme "discontinuous" jump from the minimum to the minimum plus 1, or from the maximum to the maximum minus 1. In other words, the underlying true distribution of damage (NOT pDIF) would appear to be uniform except at the endpoints.

However, from probability theory, it is known that
  • The sum of two uniform random variables with the same variance (regardless of actual minimum and maximum) follows a triangular distribution
  • The sum of two uniform random variables with different variances follows a trapezoidal distribution (so a triangular distribution can be thought of as a degenerate trapezoidal distribution)
How can I argue that the underlying random component of melee damage could follow a trapezoidal distribution?
  • Does anyone actually expect the "developers" to have done anything particularly fancy with pDIF? In many cases, random number generation is basically "sampling" from a uniform distribution, usually Unif(0,1).
  • If pDIF does follow a uniform distribution (conditional on cRatio) for some ranges of cRatio (I realize this has been shown not to be the case for certain ranges of cRatio), there could easily be another random component to introduce "jitter" into the damage calculation, which would increase the variability of damage output yet keep the mean the same. This would account for the "1.05x correction" on maximum damage I've seen bandied about from time to time.
  • So, there could be an "effective" pDIF that includes jitter.
To illustrate the plausibility of the last two points, I simulated 9,885 realizations of non-critical melee damage given 55 "base damage," with pDIF that follows Unif(1, 1.8) and a "jitter" component that follows Unif(-0.1,0.1). Here, the random components are summed together so that the end result is that the fake data are trapezoidal. I plotted the frequencies and I also drew a nonsense curve through all the data points (in blue).


Yes, this is completely fake and is not meant to demonstrate the truth of anything, but merely the plausibility that pDIF follows (or has followed) a trapezoidal distribution for some values of cRatio. I even put in a second-order polynomial trend line, which has a very high coefficient of determination, which shows that R2 cannot say anything about whether the model is even appropriate. Here, we know it is grossly inappropriate because I know what the underlying probability model is.

Here's the R code for the simulation. Data were exported to Excel. And so goes an hour of my life.

N = 9885
a = trunc(55*(runif(N,min=1,max=1.8)+runif(N,min=-.1,max=.1)))
dmg = seq(min(a),max(a))
N2 = length(dmg)
dmg.counts = numeric(N2)

for (i in 1:N2) {
dmg.counts[i] = length(a[a==(i+min(a)-1)])
}