Thursday, June 17, 2010

Does INT affect Drain accuracy?

(Correction: 06/18/2010. I meant /SCH instead of /DRK toward the end. I've gone mental...)

A while ago, I asserted that INT "seems likely" to affect the accuracy of Aspir and, by implied analogy, Drain, but I had absolutely nothing on which to base this assertion. Not quite as baseless is assuming that, since then, there has been absolutely no evidence presented anywhere to support or refute that assertion.

Some problems with getting data to show whether INT affects Drain accuracy

Why do I assume that? I'm not trying to be hater and talk shit, as ignorance about this is not on the level of, say, ignorance about the party-based, hidden latent effects of curry food items.

At least where examining the effect of INT on Drain accuracy is concerned, one problem is that if you're in a situation where you have good reason to believe Drain accuracy isn't "capped," you wouldn't want your HP to be low enough to allow you obtain the actual quantity of HP taken with Drain. (Low-level beetles and worms are not acceptable targets for examining Drain accuracy with level 75 jobs, and EM+ worms are not easily accessible... yet.)

Another problem, related to the first, is that the distribution of Drain still isn't known today, and "censored" values of HP drained don't help to provide insight into that. ("Censoring" is one way to describe the fact that Drain values reported in chat logs are based on maximum HP; any HP restored beyond your maximum HP does not count in the final chat log figure, so at best you only know at least how much you drained, not its actual value or whether your Drain was resisted.)

These are some of the problems that hamper data collection.

A way to avoid these problems?

If only there were a stationary target that didn't fight back, that could allow you to suppress your HP safely, for which Drain accuracy has the possibility not to be capped at level 75, and for which you could gather Drain data without interference from other players...

Zvahl Fortalices definitely satisfy the first condition, as they do not move. They also satisfy the second condition, as two of the fortalices deeper into Castle Zvahl Baileys (S) do not have any mobs wandering nearby, including Dark and Ice Elementals. Zvahl Fortalices definitely seemed like a promising candidate for Drain testing, so I actually set out to get some data to determine if INT has some accuracy effect.

That left the third and fourth conditions. Of course, I had no idea if it would even be possible for Drain accuracy not to be capped, but that would be part of the data collection anyway, with the hope that my Drain accuracy could be decreased enough to raise the corresponding resist rate above the (assumed) 5% resist rate floor. As for people doing skill-ups, I don't really begrudge them trying to maximize their skill-up opportunities, as this method of skill-up is liable to be "nerfed" come the June 21 version update.

Goals of data collection, some assumptions, and results

My way of determining whether INT has an effect on Drain accuracy is based on a simple two-sample comparison of the occurrence of resists, one sample based on "low" INT (71 in my case), and the other based on "high" INT (121). Again, this was based on the hope that my resist rate would not be floored (at 5%) for the low-INT case. This, in turn, is based on the assumption that the resist rate is floored at 5% and rises with decreasing magic accuracy. If the data shows the resist rate being above 5% for low INT, I conclude my Drain accuracy isn't capped for low INT. (That alone would not show that INT has an effect on accuracy; I would need the second sample under high INT as well.)

But what is considered a resist? Similar to the Aspir data collection I cited previously, it would be necessary to get some sense of the distribution of unresisted Drain values, with any low Drain values set "far enough" apart from the bulk of the data considered occurrences of a resist. This is the main assumption concerning the interpretation of the data (but a reasonable one).

Now, what about the other assumptions? I merely state some of them here because I simply was not interested in testing them, and I didn't collect enough data to test these assumptions anyway.
  • No differences among Zvahl Fortalices that could affect the results. This is a catch-all assumption concerning possible differences in magic evasion, INT, etc., but I don't think they exist (otherwise, fuck you, SE). If it could be shown that two Fortalices have two different base INT values (even if only a +1 INT difference), you would have to wonder about other possible confounders like level difference as well (can't assume these are level 75, etc.).
  • Even if there is a bonus to Drain on Fortalices, similar to a MAB bonus for elemental magic, it should still be possible to tell the difference between a resist and a non-resist. Bio II initial damage shows there is a MAB bonus, but even if there is a similar bonus for Drain (not MAB-related, of course), it shouldn't affect one's ability to distinguish between resists and non-resists.
  • The equipment bonuses (or penalties) aside from +50 INT for the "high INT" case have no effect on the accuracy of Drain. Now, obviously, I didn't put on equipment with dark magic skill or magic accuracy (or use a Dark Staff or Pluto's Staff), leaving only base attribute bonuses and penalties. Now, if you think MND and CHR actually have an effect on Drain accuracy, I'd like to hear the justification. If there are hidden accuracy effects on my equipment, that could be a problem, though.
  • Dark weather and Darksday have no effect on the accuracy of Drain. This is not really an assumption, as I didn't collect any data during Darksday or under Dark weather, but I just mention it anyway as they are potential confounders.
Now, the results. First, dot plots of the results as an initial visual impression, under low INT (top dot plot) and under high INT (below), suggest a minimum and maximum non-resisted Drain given 269 dark magic skill:


Before jumping into a discussion of the maximum and minimum Drain values, based solely on the criterion of a resist I described earlier (low values of Drain set "far enough" apart from the bulk of the data), there are 10/65 resists under low INT and 2/63 under high INT, so this data appears to provide good evidence that increasing INT increases the accuracy of Drain, especially if you think that for the 121-INT case, the resist rate was floored at 5%. I see no reason to be pedantic and report a confidence interval or p-value.

Under both low INT and high INT, the Drain maximum (unresisted) was 288 under 269 dark magic skill. It has been said that the maximum Drain and Aspir are 300 and 100, respectively, without any potency-enhancing gear (anecdotal discussion on BG), so if it can be shown that the Drain maximum is 288 under 269 skill for other mobs, you have to wonder how Drain potency actually scales with dark magic skill.

The location of the unresisted Drain minimum is less straightforward. One possibility is that it could be at 144 HP, which would be exactly half of the unresisted Drain maximum. It would be interesting if this relationship between maximum and minimum actually holds for all levels of dark magic skill (with other potency-enhancing factors presumably serving only to affect scale). One way to check this would be with with /SCH as a subjob.

And what of the relationship between the HP value of a resist and a non-resist? I actually got 29 HP under the low-INT case, and it's difficult to describe this relationship with with small samples. But small samples are enough to reach the major conclusions.

Conclusions

Based on the criterion that low values of Drain set "far enough" apart from the bulk of the observed data should be considered resists, additional INT appears to increase the accuracy of Drain. Ideally, the data collection should be repeated in an attempt to replicate this result.

Data collection could be performed using /SCH as a subjob. +50 INT (if it could be achieved) should still be able to manifest in the form of increased accuracy (provided INT does have effect), and further exploration of the relationship between dark magic skill and unresisted Drain maximum could be done, along with that between (unresisted) Drain maximum and minimum.

Monday, June 14, 2010

Probability distributions associated with WS spam

(Correction: 06/15/2010. I thought Sekkanoki lasted one minute but after reading up on it, it lasts either for one minute or until the next weapon skill, whichever comes first. So, my discussion of the consequences of TP overflow elimination now refers to a hypothetical "Sekkanoki 2.0," which would reduce the TP cost of all weapon skills to 100 TP.)

Who cares about "TP overflow"?

"TP overflow" seems to be the de rigueur term referring to any landed hits that don't contribute to spamming weapons every time 100+ TP is accumulated. TP overflow is inevitable when more than one landed hit per attack round is possible, so it's not like anyone can do much about it except attempt to minimize it by spamming WS. This absolutely does not mean it is harder to "cope" with TP overflow using a multi-hit weapon (when weapon delay is the same as a non-multi-hit alternative). Rather, slack effort means squandering the benefit of the more rapid TP gain of the multi-hit weapon.

So why care about TP overflow? One argument is that it should be "accounted for" when doing item comparisons pertaining to damage efficiency, possibly to be more accurate.

Consider, for example, Soboro Sukehiro, which is considered to average 1.9 attacks per attack round, with the probability of two attacks being .5 and that for three, .2. Given 100% hit rate and 0% DA rate, it takes 3.46553 attack rounds, on average, to be able to execute a weapon skill in six hits, with the actual average number of hits being 6.584507 (note that 6.584507/3.46533 = 1.9 attacks per round), so almost 9% of the hits occur in excess of the target number of hits.

What if somehow there was a way to allocate the TP from those excess hits toward additional weapon skills? Well, Samurai has a level 60 job ability called Sekkanoki, which limits the cost of the next weapon skill to 100 TP. This seems analogous to job abilities like Elemental Seal or Divine Seal, which lasts for 1 minute or until a spell is used, whichever comes first. But what if Sekkanoki limited the cost of all weapon skills to 100 TP while active, say, one minute? This would effectively cause a re-allocation of TP toward future weapon skills. Let's call this "Sekkanoki 2.0."

If one were under the effect of "Sekkanoki 2.0" over a very long time interval, effectively all of the TP would go toward weapon skills, and so the average number of hits approaches 6. Since the average number of attacks per round is 1.9, then the average number of attack rounds approaches 3.157894737, which seems like a fairly significant reduction in average attack rounds until you realize that the concomitant "loss" of TP damage that results from TP overflow (which is eliminated under Sekkanoki 2.0 over an infinite period of time), along with the slight loss of WS damage, offsets the benefit of increased WS frequency. (Also, the proposed Sekkanoki 2.0 lasts for 1 minute out of 5, which means that some TP overflow is inevitable for finite time periods, so it's not like Sekkanoki 2.0 has this tremendous effect.) So, the argument about accounting for TP overflow is a bit overblown (not that you shouldn't, however).

So why care about TP overflow? Since there is no Sekkanoki 2.0, which itself would be a limited tool, you can't do anything about it, so why worry about it? Maybe it's more about players wanting to appear to be "clever" about a not-very-subtle consequence of multi-hit weapons, like asserting that the probability of TP overflow for a given WS is high. (One could easily retort that for Soboro, the fraction of excess hits over total hits would be around 9%.)

But, you know, I'm all about meaningless stuff, so let's finally get into how to define the probability distribution of excess hits (that contribute to TP overflow) associated with WS spam (this would be the same as the probability distribution of the number of hits you end up with under the condition that you spam weapon skills).

Excess hits contributing to TP overflow and the corresponding probability distribution

Let E denote the number of hits in excess of those that contribute to the 100+ TP (in six hits) required to spam a WS. Let's continue with the example of Soboro. For any given attack round, the probability of n landed hits is πn, where n = 0, 1, 2, 3. These probabilities are straightforward to calculate. Not as straightforward to calculate is the probability mass function for E. An extremely tedious approach is to list all the possible combinations of attack rounds that result in 6 or more hits—the possibilities being 6, 7, or 8, which correspond to E = 0, 1, and 2, respectively. This approach requires knowing what to count (all the possible ways to get E = 0, 1, and 2), how to count (combinatorics), and knowing the closed-form expression for the sum of an infinite series, as the possibility of missing hits with non-100% hit rate means there are an infinite number of possible outcomes. (For a given combination of attack rounds leading to 100 TP, there could possibly be zero attack rounds that yield zero landed hits, one attack round that yields zero landed hits, two attack rounds that yield zero landed hits, and so on. These attack rounds are independent of those that yield hits.)

After spending more time than I care to admit, I obtained the p.m.f. of E, which is


This expression is quite unsightly, and rather useless. Not only is it useless merely because knowing the probability of TP overflow is useless, it also is useless because it refers only to the case where 6 hits are required to attain 100 TP. It requires no imagination to see that an expression for a dual-wield situation would be ghastly. It also is useless because you don't even need to knowledge of this p.m.f. to obtain the average number of hits in the process of getting to 100 TP (as I have shown repeatedly in the past). But there it is...

Again, using the Soboro example, P(E = 0) = 0.522579, P(E = 1) = 0.370335, and P(E = 2) = 0.107086, and thank goodness the probabilities sum to 1. The probability of "TP overflow" for a given WS with Soboro is almost 50%... not that you can really do anything about it. The correct response is, "who gives a shit?"

Even worse: the probability distribution of the number of attack rounds

Let R denote the (total) number of attack rounds that results in 100 TP. Again, with the Soboro example, R = 2, 3, 4, ..., and there is not much hope for an elegant formula for the probability distribution, because to obtain such a formula "by hand," one needs again to enumerate all the possible outcomes associated with each event. I only got as far as R =3 before I quit.


Again, using the Soboro example, P(R = 2) = .04, and P(R = 3) = .519. This is consistent with the average number of attack rounds being ~3, but if you already had the average number of attack rounds, why do you need the corresponding probability distribution. Useless!

A better approach for calculating these probability distributions: Markov chains

Perhaps I'll discuss this in a future entry. Aside from the fact that knowing the above probabiltiy distributions is quite useless—average weapon skill TP, average number of rounds, and average number of hits, among other things, are all easily obtained without any knowledge of these probability distributions—the Markov chain approach to obtaining these is much faster and far superior when no symbolic formulas are required. The interpretation of Markov chain output and manipulation is also much easier than it is with formulas for a specific case. It is also the only realistic way where dual-wielding is concerned, as you would have to be crazy even to consider deriving closed-form expressions for the probability distributions for that situation. It is so easy to make a mistake with a binomial or multinomial coefficient here or there, that I have to admit I didn't obtain the above expressions entirely "by hand," but with the help of Mathematica, which is quite handy for dealing with symbolic math.

Friday, June 11, 2010

How do you account for the effect of Jump?

Modeling the effect of Jump on damage rate isn't too bad provided that you invoke the following major simplifications: let both Jump and High Jump have the same amount of merit upgrades, and treat TP from Jumps as accumulating toward a weapon skill independently of TP from auto-attack. In this way, we can estimate the proportion of attack rounds that Jumps contribute to the average number of attack rounds required to accumulate 100 TP (for a weapon skill). We need this proportion to estimate the time savings from using Jumps that contribute to increasing WS frequency.

Suppose that there are 5 merits both in Jump and High Jump. This means that in a 150-second time frame, two Jumps and one High Jump can occur, for a total of three jumps. Also, do not (yet) assume that attack rounds from Jumps are equivalent to those from auto-attack in terms of multi-hit "capability" (from double attack, multi-hit weapons, etc.). It is then possible to obtain a general expression for the denominator required to obtain the respective proportions of attack rounds that auto-attack, Jump, and High Jump contribute to the average number of attack rounds to 100 TP:


The implied units for this denominator are rounds per WS (with spamming of TP after 100 TP is achieved).

The first term (factors specific to it denoted with the subscript 1) in the expression accounts for how many weapon skills from auto-attack can occur in 150 seconds when accounting for a weapon skill delay of two seconds. T1 denotes the time per attack round at 0% haste, and H denotes the haste level as an integer. E[R] in general denotes the average number of attack rounds to 100 TP, and usually, E[R1] = E[R2] = E[R3] except in the case of virtue weapons, apparently (the only reason the equality wouldn't hold because virtue weapons apparently do not work with Jumps).

The second term accounts for how many weapon skills from Jump (two Jumps in 150 seconds, remember) can occur in the previously specified 150-second time frame (necessarily a fraction), and the third term accounts for how many weapon skills from High Jump can occur in 150 seconds.

With this denominator expression, it should then be obvious how to obtain the actual proportions of attack rounds that each of auto-attack, Jump, and High Jump contribute to the average number of attack rounds to 100 TP. For example, the proportion of attack rounds that High Jump contributes to the average number of attack rounds to 100 TP is


These proportions can then be used to obtain an estimate of the adjusted average of the number of attack rounds to 100 TP accounting for Jump effects (this is a weighted average). Of course, if E[R1] = E[R2] = E[R3] = E[R], then the weighted average simplifies to E[R].

However, the adjusted average of attack rounds cannot be multiplied by a simple "time per attack round" conversion factor to get the average time to 100 TP. Recall that TP from Jumps is treated as independent of TP from auto-attack as a simplifying assumption. Instead, the aforementioned proportions must be used to obtain a weighted average of the time "per cycle" of 100 TP generated, with 2E[R2] seconds for Jump and 2E[R3] seconds for High Jump (ignoring stacking of Jump and High Jump; the units of E[R] are attack rounds "per cycle" of 100 TP generated) and E[R1]T1(100-H)/100 seconds for auto-attack.

Mechanistically, we should already recognize before doing modeling that the dominant effect of Jumps is to increase WS frequency by reducing the time required to generate 100 TP, except when T1(100-H)/100 < 2 seconds. With modeling, it is possible to estimate the reduction (both absolute and relative) in average time to generate 100 TP from Jumps. From modeling, it is also possible to account for differences in damage between auto-attack hits, Jump, and High Jump (you don't use haste equipment for Jumps, right?), but this effect is slight compared to the effect on WS frequency and will not be accounted for in future posts.

A real great katana comparison

(Correction: 06/13/2010. Additional comments are in italicized red. Incorrect statements are crossed out.)

Earlier, I blabbed about the consequences of delay associated with the use of weapon skills in terms of modeling damage output mathematically, but did not justify how much delay should be specified because I didn't know how much the following attack round (after a WS) is delayed. Fortunately, I came across this presentation of results and discussion quantifying the amount of delay that is effectively added to the attack round following the use of a job ability (or weapon skill). The results of "stacking" job abilities aside (read for yourself), it is obvious that a two-second delay for the use of a weapon skill must be accounted for, at the minimum, when attempting to model theoretical damage output. (Using other job abilities while engaging an enemy would also have an effect on damage output, but the use of weapon skills, if spammed, is the dominant factor contributing to job ability delay. Consequently, many of my previous posts, which ignored this delay, likely have led to incorrect conclusions.)

For now, though, I think it would be instructive to show how much a weapon skill delay of two seconds obviously hampers the modeling of damage output. But I don't want to waste my time doing the "before" analysis, so I base my "after" analysis based on the conditions set forth in this comparison of great katanas (covering Hagun, Soboro Sukehiro, Kurodachi, and Radennotachi). There are some problems with it, especially with the implied use of /DRG (low DA rates but not accounting for the effect of Jumps, wut). Therefore, I do not merely reuse the computed figures given but provide my own in some cases. In any case, it may help to review that comparison and mine side by side as I wish not to waste my time rehashing said conditions.

Calculating WS frequency: Zanshin is relevant for main job SAM?

The effect of Zanshin on weapon skill frequency is something I had not considered in my previous posts, and I am kind of surprised the activation rate is apparently rather high for samurai as the main job. Recall that in the October, 19, 2006 version update, "the hit rate of the extra attack [was] increased." Moreover, there is very good evidence the Zanshin activation rate can be considered 45% for main job and 25% for subjob, with the hit rate bonus the result of +35 accuracy (source). Unfortunately, it is more difficult to furnish evidence as to how Zanshin interacts with double attack for auto-attack purposes, but it seems likely that Zanshin has a lower "priority" than double attack (if double attack processes, Zanshin doesn't, and if it doesn't, Zanshin can), so I'll just run with that. This means that accounting for Zanshin doesn't really matter all that much for multi-hit weapons, but since I do it for Hagun and Radennotachi, I might as well do it for the other two.

To start off with my "after" analysis (remember I want to show the effect of weapon skill delay not previously considered on an analysis that incorrectly ignores it), going back to the "before" analysis I cited previously, I should first point out that pDIF is apparently ignored in favor of a bogus assumption of a "baseline" 35:65 ratio of melee damage to WS damage for Hagun.

Since we are talking about theoretical damage output, it is nonsense to assume such a ratio. The baseline assumption is bogus, not that 35:65 may be observed in practice. If 35:65 is observed, surely average auto-attack damage and average WS damage are also observed (from parser output)! Use those values instead to back-calculate an "average" pDIF for both auto-attack and WS damage that should be fixed across all great katanas. The differences in WS frequency and weapon base damage will then account for the differences in the ratio of melee damage to WS damage, holding pDIF constant.

Anyway, I will return to the pDIF issue later. After accounting for the effect of Zanshin, I obtain the following averages for attack rounds from WS use to 100 TP, auto-attack hits in the process of getting to 100 TP after WS use, and the "effective" hit rate (landed hits per attack round), which encompasses the effects of accuracy, double attack, and Zanshin.

Weapon
Average no.
of rounds
Average no.
of hits
Effective hit rate
Hagun
5.059455.107421.00948
Soboro Sukehiro
3.11051
5.58467
1.79542
Kurodachi
3.99375
5.326421.33369
Radennotachi
5.05945
5.10742
1.00948

I am aware of the apparent absence of Brutal Earring (5% DA) for Soboro (but why use a Pole Grip then, implied with the stated 2% DA?), replaced by a mysterious source of accuracy +5, and accounted for those differences. I gave the benefit of the doubt, so to speak, with Soboro (94% hit rate after accuracy +5), even though it could easily be argued that, across all merit mobs encountered, the average hit rate could actually be closer to 93.5%.

My effective hit rate figures agree with the previous analysis more or less, but I do not compute effective hit rate directly. Instead, I compute it, as a kind of check on my calculations, after computing the average number of attack rounds and average number of hits (example: 5.10742/5.05945 = 1.00948) to make sure I didn't make any errors calculating the average number of attack rounds.

As always, the average number of rounds can be converted to the average time to accumulate 100 TP, but now the time between weapon skills must also account for the two-second weapon skill delay discussed previously. (This will be done at the end of the post.)

Accounting for average TP for the use of Tachi: Gekko

The previous analysis assumes maximum fSTR for each of the weapons (16, 12, 15, and 17 for Hagun, Soboro, Kurodachi, and Radennotachi, respectively), which would appear to be reasonable given the implied high STR modifier bonus used for Tachi: Gekko (152*.75*.83 = 94.62, which is close to the given 94). As mentioned previously, pDIF is completely ignored, but based on the attack bonus of Tachi: Gekko, it is reasonable to assume an average pDIF of 2.3 (based on a symmetric pDIF distribution between 1.9 and 2.7).

The only thing left is calculating the fTP bonus of the first hit for Tachi: Gekko, which requires calculation of average TP for each weapon when a one-hit weapon skill is used, accounting for double attack. This, in turn, requires knowledge of the probability distribution of TP return from a one-hit WS and the corresponding TP values, which is the same regardless of weapon.

This would seem straightforward except for the observation of 2-TP return with one-hit weapon skills (source), which would suggest that for weapon skills, Zanshin can occur on the first hit independent of the double attack (Zanshin still can't occur for the double attack hit, presumably). The presence of Zanshin effectively "reallocates" the probability of missing the first hit (and losing the full TP return of 16.7), which is 5% most likely, so ignoring the Zanshin effect for a one-hit weapon skill results in negligible error for TP return (but not necessarily WS damage).

Weapon
Average TP per WS
(my calculation)
fTP bonus of 1st hit
(with Gorget effect)
Hagun
101.278061.9829879
Soboro Sukehiro
109.24816
1.6914005
Kurodachi
104.935331.6779229
Radennotachi
101.27806
1.6664939

Note that average TP shouldn't be truncated because these averages are themselves based on the actual truncated TP figures to begin with (assumed 16.7 TP per main WS hit and auto-attack hits and 1.4 TP for off-hand WS hit).

Accounting for average Tachi: Gekko damage: ignore Zanshin?

Given 91% hit rate for any double attack hits (7% DA rate) for Tachi: Gekko (95% otherwise), the average number of hits per weapon skill is .95 + (.91)(.07) = 1.0137. Accounting for the 45% Zanshin rate, this average rises to 1.035075, of which .95 still corresponds to the first hit (which receives the fTP bonus), so 0.085075 of the hits in the average WS have an fTP = 1. The effect of Zanshin is, therefore, like adding 2.345% DA, which, for the purposes of Tachi: Gekko, constitutes approximately a 1.1-1.3% increase in average WS damage. (This is given the conditions stated in the "before" analysis). Whether or not this is accounted for (I will account for it), the effect of Zanshin very slightly "favors" weapons with worse WS "secondary" hit damage (compared to other factors), so it can be ignored for convenience.

A "fatal" flaw: consequences of the effect of haste with weapon skill delay

Because weapon skill delay, which is a fixed value (consider it two seconds), exists, the relative benefit of haste (or other forms of delay reduction) is higher for weapons with lower weapon-skill frequency compared to weapons with higher weapon-skill frequency. It follows that a weapon with higher weapon skill frequency CAN actually be "worse," on average, than a weapon with lower weapon skill frequency depending on the level of haste!

One way to think of this is to consider an arbitrary time frame during which weapon skills occur. The time associated with the WS frequency might be reduced with haste, but there is always an absolute weapon skill delay tacked on. Even if haste goes to 100% (meaning the time associated with WS frequency goes to 0) and you still decide to use WS for some reason, the sum of the absolute weapon skill delay for the weapon with higher WS frequency will be higher than equal to that for the weapon with lower WS frequency (WS frequency is rendered irrelevant if it takes zero time to build TP toward a WS), so the weapon with higher WS damage wins out in terms of WS damage output.

A "practical" consequence is that for "zerging" situations where maximum haste is involved, low-damage, multi-hit weapons (on average) can be worse than standard weapons. Similarly, multi-hit weapons may not be that good for meriting situations.

The "fatal flaw" with the "before" analysis is the unstated assumption that the haste level doesn't matter across weapons, so that the "pecking order" of great katanas always holds. Because weapon skill delay is not accounted for, the analysis does not hew to what is experienced in practice.

Repeat the analysis instead with ~65% haste (Hasso, Haste spell, double March, 20% equipment haste) along with the weapon skill delay of two seconds. The following figures are the result of a "per weapon skill" perspective, using average auto-attack pDIF 1.15 and average WS pDIF of 2.3. (Overwhelm 5/5 also used.)

Weapon
Avg. TP dmg
Avg. WS dmg
Time per WS
Dmg/sec
TP:WS dmg
Hagun
510.99763910.7053192
15.281 s
93.04
36:64
Soboro Sukehiro
333.96349
619.4514917
10.165 s
93.7935:65
Kurodachi
490.03068
744.1553438
12.810 s
96.35
397:603
Radennotachi
593.22713
835.6201688
15.281 s
93.50
415:585

Given 65% haste, relative to Hagun, Kurodachi is about (96.3474/93.0370 - 1)100% = 3.56% more efficient, and Soboro, about (93.7931/93.0370 - 1)100% = 0.81% more efficient. Radennotachi is about 0.5% more efficient. This jibes with the observation that Soboro is not really any better than Hagun in a typical merit situation.

Now, what happens given 80% haste?

Weapon
Avg. TP dmg
Avg. WS dmg
Time per WS
Dmg/sec
TP:WS dmg
Hagun
510.99763910.7053192
9.589 s
148.2636:64
Soboro Sukehiro
333.96349
619.4514917
6.666 s
143.03
35:65
Kurodachi
490.03068
744.1553438
8.177 s
150.93
397:603
Radennotachi
593.22713
835.6201688
9.589 s
149.01
415:585

Obviously, the damage figures (other than rate of damage) shouldn't change with haste. As they are fixed, changes in relative efficiency calculations (relative to 65% haste) involve only changes in time per WS (where applicable). The effect of 15% more haste benefits Hagun relatively more than it does Soboro because of the presence of the fixed two-second weapon skill delay. The result here shows that Hagun is more efficient than Soboro in a max-haste situation when spamming WS, and you should be. 910 damage, on average, in exchange for 2 seconds is better than 511 damage, on average, over 7.589 seconds.

(Correction: 06/13/2010) Incidentally, given 9% DA (the stated condition), Kurodachi is still better than Hagun even with maximum haste, so it is just better barring situations where WS damage is the predominant form of damage and WS frequency is an irrelevant consideration but as DA increases, Hagun eventually becomes better than Kurodachi. This should make sense (but even I overlooked this...) because the "full" benefit of a DA increase is not realized with multi-hit weapons such as Kurodachi, and definitely not with Soboro Sukehiro.

Conclusion

Weapon skill delay, which exists and can be considered to be two seconds, should be considered when doing a theoretical comparison of things related to doing damage.

A major consequence of weapon skill delay is that, as haste increases, weapons with lower WS frequency benefit relatively more than weapons with higher WS frequency. This affects the "correct" choice of weapon for situations where high levels of haste are achieved. For example, even though Soboro Sukehiro may be better than Hagun at low levels of haste, it is inferior at high levels of haste (on average, since there is some inherent variability of WS frequency associated with multi-hit weapons).

(Correction: 06/13/2010) However, it can be shown that Kurodachi is superior to Hagun when WS frequency is a relevant factor (e.g., not relying only on Meditate to generate TP). "Actually better" in theory, however, is contingent on how much base DA is present.

The effects of Zanshin on WS frequency, WS damage (fTP bonus and Zanshin hits), and TP return can be quantified. While the effects of Zanshin given low hit rate were not discussed, the effect of Zanshin can "safely" be ignored for relative comparisons given high hit rates.

Tuesday, May 25, 2010

What's the proc rate for virtue weapons? How do you know?

One "line" of evidence: checking with Justice Sword

Taken at face value, the estimate 555/1000 indicates the "occasionally attacks twice" (OAT) rate of Justice Sword is significantly higher than 50% and could be considered 55% (source). But is the OAT property the same for all so-called "virtue weapons"?

Another line of evidence: checking with Fortitude Axe... and WAR

The rest of this post discusses how to estimate the OAT rate of Fortitude Axe in the presence of the double attack trait from WAR. But first, I needed a good idea about how Fortitude Axe OAT actually interacts with the DA trait. In the past, I blabbed a lot about how Fortitude Axe might interact with double attack, but my "conclusion" was based on very weak evidence. After collecting some more count data with kparser under 12% DA (source), which ruled out my previous weak hypotheses about the DA/OAT interaction, I got a better idea about how to explain these results (assuming kparser was working correctly...).

It appears (not exactly "proof") that OAT can process on both the normal hit (which is guaranteed to occur for a given attack round, if not actually land) as well as the hit from the possible DA proc (with a major caveat to be discussed soon). More specifically, the hit from the DA proc occurs independently of whether an OAT proc occurs. (Not probabilistically, of course, but "mechanistically.")

It is worth noting that conceptually the order of DA and OAT could easily be reversed, such that the hit from the OAT proc occurs independently of whether a DA proc occurs, but I will just say the resulting probability calculations are not supported by the data when OAT is mechanistically independent of DA.

Anyway, one way to show pictorially all "possible" outcomes where DA and OAT can interact is with the following "tree":


There are six hypothetically "distinct" outcomes, but it is very inconvenient to monitor the equipment menu for virtue stone expenditure. More important, though, is the fact that Fortitude Axe cannot quadruple attack, so the case of expending two virtue stones is impossible. (This makes sense, noting that triple attacks are impossible with zero DA rate.)


So what "happens" to this 2-virtue stone attack round that is impossible? It appears that even if a DA proc occurs, only one virtue stone can be expended anyway, so the "tree" simplifies further:


The resulting probability model of the number of hits in an attack round (ignoring the distinction between hits and misses) is specified as follows. Let X denote the number of hits in a given attack round, d the probability of a double attack proc, and π the probability of an OAT proc. Then,

Now that we have a reasonable probability model describing the interaction between double attack and the OAT property of Fortitude Axe ("reasonable" based on chi-square goodness of fit to the data given 12% DA rate and posited virtue weapon proc rates of 50% and 55%), we can now estimate the OAT rate. Proceeding with maximum likelihood estimation is not really necessary when an obvious unbiased estimator can be based off the observed number of single hits (denoted as X1) in n attack rounds:


It follows that the unbiased estimator is


with variance


Note that when d = 0, the variance reduces to that for the estimator for a simple binomial proportion (marginal in the context of the multinomial distribution). (Note to self: from simulation, this estimator is only very slightly less efficient, from an MSE standpoint, than the MLE, which I would bet is UMVUE even if an analytical expression for the MLE and the CRLB is annoying to obtain.)

The estimated proportion of single hits (per attack round) is 1 - 595/1425/.88 = .5255183, with corresponding 95% confidence interval (.4964218, .5546148). Given the specified probability model (which cannot be "proven" to be true at this time) and the data, it is not possible to conclude that the OAT rate of Fortitude Axe is either 50% or 55% (both are plausible given the confidence interval), unfortunately. But it should be possible to rule out one or the other with further data collection (with the hope that the probability model is correct), using the estimator specified above.

A third way: Faith Baghnakhs

Among all virtue weapons, it would be fastest to determine the OAT rate of Faith Baghnakhs by counting the number of triple attacks and quadruple attacks. It would be easier to do this on ninja because you wouldn't have to pay attention to kick attacks (because you want to use a parser instead of counting manually). If the OAT rate for Faith Baghnakhs can be shown to be 55%, that, along with the observed proc rate for Justice Sword, could be used as evidence for a common OAT rate of 55% across all virtue weapons.

Monday, May 17, 2010

A hierarchy of great axes?

This is a rehash of a previous post comparing Bonesplitter and the good Luchtaine, two "Magian" great axes, to that old standby Perdu Voulge and Fortitude Axe, the presumptive weapon of choice for Campaign (even though Waltz recast ends up being the rate-limiting factor for curing yourself), but new evidence, both for Fortitude Axe (see first relevant BG post and second relevant BG post for details that I won't go over here) and Luchtaine (to be discussed later, perhaps), show that I underrated Fortitude slightly and overrated Luchtaine significantly.

In particular, evidence indicates Luchtaine behaves similarly to Joyeuse such that regular DA and Magian OAT are "directionally" exclusive, which is different than mutually exclusive. Suppose that the DA rate were 20%. Then, mutually exclusive would mean P(OAT) = .40, P(DA) = .20, and P(OAT and DA) = 0. On the other hand, directionally exclusive would mean that either P(OAT|not DA) = .40 and P(DA) = .20 OR P(DA| not OAT) = .20 and P(OAT) = .40. Consequently, given 20% DA and 40% OAT rate, the effective DA rate would be .20 + .80*.40 = .40+.60*.20 = .52.

Also, I decided to repeat the previous analysis using Raging Rush. Even if Raging Rush's three base hits (in other words, those not arising from double attack) are the only ones that have a chance to be critical hits, it's still generally better than King's Justice. One consequence: because RR's STR modifier is lower than KJ's, the relative difference in damage between a Perdu RR and Fortitude RR is more than that between a Perdu KJ and Fortitude KJ, so the relative difference between Perdu and Fortitude "overall" would be less with RR than KJ "all other things being equal."

I find it is worth including Rune Chopper in the discussion, too, along with Hephaestus with STR +4 and attack +15 as a basis of my pontificating about what kind of effort is warranted to get "good enough" (not the most). Since I have 19% haste normally, I will use that as a haste baseline before Rune Chopper, so the full haste bonus of Rune Chopper is not fully realized. On the other hand, I will also consider having Rune Chopper with only 1 MP refresh such that the latent is active one out of every two rounds (as it appears to be anyway). I will also consider the situation of having a "typical" double March (~20% haste with March +2 instrument and 8/8 merits in both wind and singing skill), Haste spell (~15%), and Hasso (~10%).

I will also account for the concept of time delay between the initiation of a weapon skill and the start of the next attack round, as it apparently is fundamental to the game and not associated with human reaction time or laziness (not that I really noticed or cared), kind of like how the delay associated with Curing Waltz screws up Drain Samba actually working properly (something that is easy to notice and that I find very annoying). This could also be considered the time delay associated with execution of a weapon skill that must elapse before the start of the following auto-attack round, or "WS delay" for short. It's something worth considering because this delay is unavoidable, but since I don't know what is actually the so-called WS delay, I will do this comparison for 0, 1, 2, 3, and 4 second delays.

Finally, I find it really unnecessary to go into excruciating detail about what goes in the calculations, so I will just report something I call "relative efficiency" ratios relative to the baseline of Perdu Voulge, which are merely ratios of damage rates. In the end, one should focus only on the gross differences, rather than whether something is really 2.15% more as opposed to 2.2% more, for example.

Relative efficiency of great axes relative to Perdu Voulge (in terms of damage rate)

Weapon
No WS delay
1s delay
2s delay
3s delay
4s delay
Rune Chopper
(latent active always)
1.0971.082
1.0701.0591.049
Fortitude Axe
1.0371.008
0.984
0.9640.947
Hephaestus
(6 hits to 100 TP)
1.0321.029
1.0261.0231.021
Bonesplitter
1.0171.017
1.017
1.017
1.017
Perdu Voulge
11
111
Rune Chopper
(latent 1/2 active)
0.997
0.991
0.986
0.9810.977
Luchtaine
0.9850.972
0.9610.9520.944
Hephaestus
(7 hits to 100 TP)
0.9530.961
0.968
0.974
0.980

Again, these ratios are based on 19% equipment haste before Rune Chopper, along with double March (~20% march), Hasso, and Haste spell.

The ratios under the hypothetical situation with zero WS delay can represent the "intrinsic" relative efficiency of weapons that have a higher WS frequency that Perdu Voulge (notice that Bonesplitter has the same relative efficiency regardless of WS delay because it has the same WS frequency as Perdu), but intrinsic doesn't mean actual or true. The higher the WS delay, the more disproportionately affected are weapons with higher WS frequency compared to Perdu.

(Note that the concept of WS delay can be generalized to job abilities that interrupt or postpone attack rounds, but I did not account for that here.)

Implications for Fortitude Axe: wonder why Fortitude Axe doesn't actually appear to be better than Perdu Voulge in practice? WS delay could explain it. In particular, if you plan to use Fortitude Axe in a maximum haste situation (~80% haste), spamming weapon skills might be relatively counter-productive (for Fortitude compared to Perdu) because of WS delay, but WS frequency is pretty much the only benefit of using Fortitude Axe (aside from TP gain without using WS), so why not just use a high-damage great axe (Perdu or even Berserker's Axe)? Using Fortitude Axe for a zerg basically means having hope that you get more hits per round in a small time frame compared to the long-run average, e.g., stringing together several 3-attack rounds. Having 80% haste is usually the decisive factor in a max-haste zerg because you probably have max attack and accuracy as well. Maybe if you had a BLM land Choke and got some STR etudes... things you could do to compensate for the low base damage.

Implications for Rune Chopper: On the other hand, Rune Chopper with latent always active (if you somehow manage to achieve this; not a trivial thing) is still substantially better than Perdu even with significant WS delay (4 seconds), and this is under the situation where the 9% haste bonus isn't fully realized (it could be if you switched out other haste equipment to increase other damage-related factors), albeit under the double March/Hasso/Haste spell situation. On the other hand, Rune Chopper with only 1 MP refresh is rather pointless. If you had an Ares Cuirass lying around and RDM accommodating you, it would be good.

Implications for Luchtaine and other Magian great axes: SE really needs to allow Luchtaine to attack 3 times or even 4 times in the future or increase the base damage dramatically. At least there is hope that SE might do this later, whereas with Fortitude Axe, SE will never allow 4 attacks per round. You don't get much out of the others (as the final forms currently are) considering the time investment required, compared to spending IS on a Perdu Voulge. Hephaestus 6-hit is not terribly reasonable because of the 29 store TP requirement alone.

Didn't I say something about a hierarchy? Stick with Perdu Voulge in general...

Tuesday, May 11, 2010

How to check if a Teiwaz has superior accuracy to Terra's Staff

Obviously, I am talking about the Teiwaz with elemental affinity: magic accuracy +3, not so much that I am talking about the earth-aligned Teiwaz.

It is thought that earth affinity: magic accuracy +1 is equivalent to the accuracy bonus of Earth Staff, and earth affinity: magic accuracy +2 equivalent to the accuracy bonus of Terra's Staff. If you ever read this blog, you would know that elemental NQ staves are considered to have +20 magic accuracy for the specified element, and HQ staves, +30 magic accuracy. It is postulated that elemental affinity: magic accuracy +3 corresponds to +40 magic accuracy.

Without discussing the evidence underlying the following experiment to check whether the earth Teiwaz is superior to Terra's Staff, I will describe a superiority "trial" involving relatively few casts.

Location: Alzadaal Undersea Ruins (Nyzul Isle Staging Point)
Target monster: Level 78 Qiqirn Poulterer (ranger)
Spell to cast: Stone (I)
How many casts: 100.
What to count: number of non-resisted Stone I, number of half-resisted Stone I, number of quarter-resisted Stone I, number of eighth-resisted Stone I (should be easy to identify from the magnitude of damage)

How to identify the level 78 Qiqirn Poulterer: one way to check you have found the correct level Qiqirn is to set your accuracy score to 263 and use the "check" function to find the right Qiqirn. One way to achieve this is to equip a weapon type for which you have 230 combat skill (example: BLM with max club skill). 230 combat skill corresponds to 227 accuracy. Suppose you also have 62 DEX. For a one-handed weapon (club), this means you have +31 accuracy. Then equip +5 accuracy worth of equipment (example: Chivalrous Chain) to achieve a total accuracy score of 263. Level 77 and 76 Poulterers will check "low evasion," while the level 78 Poulterer will give no evasion message. Incidentally, this implies the level 78 Poulterer has at least 293 total evasion. You can confirm the level after killing the Poulterer by noting EXP yield (200 base EXP for level 78, 230 given 15% Sanction bonus).

Total magic accuracy for this experiment: it is known reasonably well (I will not cite evidence at this time) that having 65 INT, 290 elemental magic skill, +5 magic accuracy (from equipment), and no elemental staff corresponds to having about 55% magic accuracy rate for the Stone I spell. (The level 78 Qiqirn Poulterer has 65 INT.) If your elemental magic skill is higher or lower (say 292), make the appropriate adjustments to INT and/or magic accuracy. Here, +/-1 INT is considered +/-1% magic accuracy rate (up to a point), and +/-1 magic accuracy (from equipment) is considered +/-1% magic accuracy rate, too.

Given the above, equipping a HQ staff like Terra's brings your magic accuracy up to ~85%. A "quickie" trial I ran gave 57/71 non-resisted Stone I, strong evidence of uncapped magic accuracy rate. As postulated previously, equipping a Teiwaz with earth affinity: magic accuracy +3 could bring your magic accuracy up to ~95% (the maximum rate).

Why 100 casts of Stone I on a level 78 Qiqirn Poulterer? 100 is an arbitrary figure as I am too lazy to do a power calculation, but since we know a priori that Terra's Staff doesn't even give a capped magic accuracy rate (and it shouldn't since I said it would be ~85%), it will be very easy to show that, if the Teiwaz earth affinity: magic accuracy +3 really has +40 magic accuracy, the observed data will indicate a capped magic accuracy rate. Of course, if the accuracy bonus were higher than +40, this test wouldn't be able to show that, but the in-game constraints described by the current magic accuracy "model" and a desire for a "minimal" sample size (the further away from 50%, the smaller the standard error) led to the above experimental conditions.

Considerations: avoid Qiqirn Goldsmith links. To minimize damage from ranged attacks, it is preferable to be RDM. If not, have someone spam heal you while you hammer out 100 casts in short order. It really doesn't take that long.

Questions and desired clarifications about experimental conditions may be fielded in the comments, if anyone actually gives a shit.

Credit to pchan on BG for previous work on Qiqirn Poulterers that allows for a fairly straightforward and not-onerous experiment.