|
|
On December 06 2013 05:54 SKC wrote:Show nested quote +On December 06 2013 05:53 hariooo wrote:On December 06 2013 05:47 SKC wrote:On December 06 2013 05:45 hariooo wrote:On December 06 2013 05:39 SKC wrote:On December 06 2013 05:36 hariooo wrote:On December 06 2013 05:33 Klonere wrote:On December 06 2013 05:29 crms wrote:I don't care how accurate, inaccurate, bullshit or truthful DSR ends up being, the entire site and all the drama was worth it just for this: http://i.imgur.com/VzbcciG.pngTrolling friends about DSR has been more fun this week than actually playing DOTA. its good for this as well Bulldog queues pubs with Sheever a lot. This particular example isn't too surprising. The fact he queues with someone doesn't really change how good of a player he is. It means the average MMR of all players in those games are going to be lower than if he was solo q'ing. DSR is to some significant extent (from the creator's mouth) based off that value. I can't imagine Bulldog going EE-levels of tryhard in those kinds of games either, so it shouldn't shock anyone that he might have a lower DSR compared to pubstars. 2GD is a pubstar? Just because Admiral Bulldogs MMR is lower than it could be doesn't mean he should be that low. They also don't queue exclusivelly with each other. 2GD has even talked about how his games are very diferent when he doesn't stack with Draskyl. His MMR isn't that high at all. Right but MMR and DSR are different measures and not directly comparable as far as I know. If 2GD frequently stacks with people of higher MMR than him and Bulldog frequently stacks with people of lower MMR than him, this is an explanation for the difference in DSR's. Oh, sure, that could definatelly be the reason why they are diferent. Which is also why Dota Skill Rating is a shit Skill Rating if it doesn't rate skill. I don't think anyone is arguing MMR and DSR are not very diferent from each other.
At least this is more debatable. Think about if LeBron James only played street ball with his friends. He's undoubtedly really good, but it's near impossible to determine his full skill potential at pro level using any rating system.
It sounds like DSR is only valid for solo q'd games. This can be a weakness of course, or you could just consider that its results are only valid under certain circumstances.
|
Or you could consider that it's results are hilariously invalid in all circumstances.
|
On December 06 2013 06:02 hariooo wrote:Show nested quote +On December 06 2013 05:54 SKC wrote:On December 06 2013 05:53 hariooo wrote:On December 06 2013 05:47 SKC wrote:On December 06 2013 05:45 hariooo wrote:On December 06 2013 05:39 SKC wrote:On December 06 2013 05:36 hariooo wrote:Bulldog queues pubs with Sheever a lot. This particular example isn't too surprising. The fact he queues with someone doesn't really change how good of a player he is. It means the average MMR of all players in those games are going to be lower than if he was solo q'ing. DSR is to some significant extent (from the creator's mouth) based off that value. I can't imagine Bulldog going EE-levels of tryhard in those kinds of games either, so it shouldn't shock anyone that he might have a lower DSR compared to pubstars. 2GD is a pubstar? Just because Admiral Bulldogs MMR is lower than it could be doesn't mean he should be that low. They also don't queue exclusivelly with each other. 2GD has even talked about how his games are very diferent when he doesn't stack with Draskyl. His MMR isn't that high at all. Right but MMR and DSR are different measures and not directly comparable as far as I know. If 2GD frequently stacks with people of higher MMR than him and Bulldog frequently stacks with people of lower MMR than him, this is an explanation for the difference in DSR's. Oh, sure, that could definatelly be the reason why they are diferent. Which is also why Dota Skill Rating is a shit Skill Rating if it doesn't rate skill. I don't think anyone is arguing MMR and DSR are not very diferent from each other. At least this is more debatable. Think about if LeBron James only played street ball with his friends. He's undoubtedly really good, but it's near impossible to determine his full skill potential at pro level using any rating system. It sounds like DSR is only valid for solo q'd games. This can be a weakness of course, or you could just consider that its results are only valid under certain circumstances. I don't see how useful a stat that only works in circumstances that don't exist can be. Almost noone plays solo queue exclusivelly, and the website is definatelly not making the distinction. It just means the currest rating is useless.
And that's giving the benefit of the doubt that it would actually work in this imaginary scenario.
|
On December 06 2013 05:55 Sn0_Man wrote:
Doing something like swapping your carry to safety as venge literally lowers your DSR. HOW STUPID IS THAT?
all systems have issues, even the mightly ELO. I mean you can be 30-0-30 and lose and that LOWERS YOUR MMR/ELO, HOW STUPID IS THAT??! (sarcasm)
The DSR method while certainly riddled with issues is still 'something'. Does it work perfectly? Nope. Are there flaws? Yep. But for all the trolling and complaining it's been more or less about where I'd expect it to be for most people I play with. My friends I expected to be close to me have been close, worse than me have been worse, and better than me better. Sure it's anecdotal evidence but if you take it with a grain of salt, it hasn't been so bad.
The main point and underlying conclusion of all the DSR discussion to me has been that ratings and the desire for ratings is obviously super hyped in the community and I wish Valve would just step in and deal with it.
|
On December 06 2013 06:06 crms wrote:Show nested quote +On December 06 2013 05:55 Sn0_Man wrote:
Doing something like swapping your carry to safety as venge literally lowers your DSR. HOW STUPID IS THAT? all systems have issues, even the mightly ELO. I mean you can be 30-0-30 and lose and that LOWERS YOUR MMR/ELO, HOW STUPID IS THAT??! (sarcasm) The DSR method while certainly riddled with issues is still 'something'. Does it work perfectly? Nope. Are there flaws? Yep. But for all the trolling and complaining it's been more or less about where I'd expect it to be for most people I play with. My friends I expected to be close to me have been close, worse than me have been worse, and better than me better. Sure it's anecdotal evidence but if you take it with a grain of salt, it hasn't been so bad. The main point and underlying conclusion of all the DSR discussion to me has been that ratings and the desire for ratings is obviously super hyped in the community and I wish Valve would just step in and deal with it. The part that you are missing is that assuming the ELO of everybody else in the game is accurate, you are expected in that game to do something better than you did if you wish to win. I mean, if I play with 4 bots versus a team of 5, and the ELO is statistically balanced, I'm probably expected to go 100-0 or else the bots might just feed the game away. Extreme example but you get the idea.
Regardless, yes there is WIDESPREAD community demand for a rating system. DBR was actually quite accurate. DSR is actually just a joke and the people who are defending it are truly hilarious.
|
On December 06 2013 06:00 Milkis wrote:Show nested quote +On December 06 2013 05:58 devilesk wrote:On December 06 2013 05:50 Milkis wrote: I don't really want to go through all that PHP but it seems like the general idea was to train a bunch of data to get stuff like kda/xpm/gpm and such for every hero and judged you by the performance of those statistics per hero. That would be my guess. This is why he can get away with "20 games" for that reason.
I think such a methodology can work very well to rate players but I don't think how he's accomplishing is the best way ever. There's a lot of issues with such a system too that I wonder if it was addressed (I don't really want to read that code). But it's really not difficult to think which ones work or not.
One more thing is that every rating system has anomalies -- I mean, not everyone tryhard ladders every game and I think things like that affect your rating probably. No, he might train with all that data, but ultimately the question he's trying to solve with his model is whether a game is "pro-level" or "normal". From his training, he can "graph" a surface that represents the baseline between pro and normal. Anything below the surface is normal and anything above it is pro. He's using the distance from that surface as the metric for generating his rating, which is I believe a point of contention for people knowledgeable about SVM (I am not), because it seems like the model's simple purpose is just to make that binary decision. Yeah I'm definitely not defending how he implemented it -- I'd have approached it very differently. There's a lot to criticize on the implementation, but the idea itself isn't without merit IMO Like I think if he just looked at stuff like CS, game length, etc, to control for that he'd have a much stronger rating. But I agree that doing that based on just one GPM/XPM/KDA estimate per a single hero isn't the best idea ever.
My problem isn't with what variables he's taking into account. My problem is with the question his model is trying to answer and how I don't see how it's even good for basing a rating system off of.
So he's trying to classify games as in pro bracket or non pro bracket. Suppose he can actually do this perfectly, taking into account every stat imaginable. So now I can tell you that 0 out of your 500 games were played at a "pro level". What good is that metric, especially if 90% of everyone else has the same result? Even if the two brackets were distributed evenly, so then 250 out of your 500 games were played at a "pro level" how do you create a rating out of that?
Since Valve MM matches people based on a continous scale, there's more than just "pro" and "normal" bracket. I can say things like this match in the 99th percentile of matches based on the average Elo in the game, or it was played at the 5th percentile. The only thing you can do with the metric this guy is measuring is say, this match was played in "pro bracket" or this match was played in "normal bracket"
|
On December 06 2013 06:05 SKC wrote:Show nested quote +On December 06 2013 06:02 hariooo wrote:On December 06 2013 05:54 SKC wrote:On December 06 2013 05:53 hariooo wrote:On December 06 2013 05:47 SKC wrote:On December 06 2013 05:45 hariooo wrote:On December 06 2013 05:39 SKC wrote:On December 06 2013 05:36 hariooo wrote:Bulldog queues pubs with Sheever a lot. This particular example isn't too surprising. The fact he queues with someone doesn't really change how good of a player he is. It means the average MMR of all players in those games are going to be lower than if he was solo q'ing. DSR is to some significant extent (from the creator's mouth) based off that value. I can't imagine Bulldog going EE-levels of tryhard in those kinds of games either, so it shouldn't shock anyone that he might have a lower DSR compared to pubstars. 2GD is a pubstar? Just because Admiral Bulldogs MMR is lower than it could be doesn't mean he should be that low. They also don't queue exclusivelly with each other. 2GD has even talked about how his games are very diferent when he doesn't stack with Draskyl. His MMR isn't that high at all. Right but MMR and DSR are different measures and not directly comparable as far as I know. If 2GD frequently stacks with people of higher MMR than him and Bulldog frequently stacks with people of lower MMR than him, this is an explanation for the difference in DSR's. Oh, sure, that could definatelly be the reason why they are diferent. Which is also why Dota Skill Rating is a shit Skill Rating if it doesn't rate skill. I don't think anyone is arguing MMR and DSR are not very diferent from each other. At least this is more debatable. Think about if LeBron James only played street ball with his friends. He's undoubtedly really good, but it's near impossible to determine his full skill potential at pro level using any rating system. It sounds like DSR is only valid for solo q'd games. This can be a weakness of course, or you could just consider that its results are only valid under certain circumstances. I don't see how useful a stat that only works in circumstances that don't exist can be. Almost noone plays solo queue exclusivelly, and the website is definatelly not making the distinction. It just means the currest rating is useless. And that's giving the benefit of the doubt that it would actually work in this imaginary scenario.
There's been enough evidence (DSR predicting wins at high percent is the big one) that I believe DSR works in the optimal scenario (solo Q for last 20 games against other solo q'ers).
The fact that it loses a lot of its accuracy under suboptimal conditions is certainly a disadvantage. It's the most legitimate complaint I've heard against it all thread.
|
On December 06 2013 06:06 crms wrote:Show nested quote +On December 06 2013 05:55 Sn0_Man wrote:
Doing something like swapping your carry to safety as venge literally lowers your DSR. HOW STUPID IS THAT? all systems have issues, even the mightly ELO. I mean you can be 30-0-30 and lose and that LOWERS YOUR MMR/ELO, HOW STUPID IS THAT??! (sarcasm) The DSR method while certainly riddled with issues is still 'something'. Does it work perfectly? Nope. Are there flaws? Yep. But for all the trolling and complaining it's been more or less about where I'd expect it to be for most people I play with. My friends I expected to be close to me have been close, worse than me have been worse, and better than me better. Sure it's anecdotal evidence but if you take it with a grain of salt, it hasn't been so bad. The main point and underlying conclusion of all the DSR discussion to me has been that ratings and the desire for ratings is obviously super hyped in the community and I wish Valve would just step in and deal with it. We all know I'm better than you Jeff.
|
On December 06 2013 06:11 Comeh wrote:Show nested quote +On December 06 2013 06:06 crms wrote:On December 06 2013 05:55 Sn0_Man wrote:
Doing something like swapping your carry to safety as venge literally lowers your DSR. HOW STUPID IS THAT? all systems have issues, even the mightly ELO. I mean you can be 30-0-30 and lose and that LOWERS YOUR MMR/ELO, HOW STUPID IS THAT??! (sarcasm) The DSR method while certainly riddled with issues is still 'something'. Does it work perfectly? Nope. Are there flaws? Yep. But for all the trolling and complaining it's been more or less about where I'd expect it to be for most people I play with. My friends I expected to be close to me have been close, worse than me have been worse, and better than me better. Sure it's anecdotal evidence but if you take it with a grain of salt, it hasn't been so bad. The main point and underlying conclusion of all the DSR discussion to me has been that ratings and the desire for ratings is obviously super hyped in the community and I wish Valve would just step in and deal with it. We all know I'm better than you Jeff. nice dsr jadelord.
only 1 c-god in irc bro and he ain't from chicago.
(iloveyoucomeh)
|
On December 06 2013 06:13 crms wrote:Show nested quote +On December 06 2013 06:11 Comeh wrote:On December 06 2013 06:06 crms wrote:On December 06 2013 05:55 Sn0_Man wrote:
Doing something like swapping your carry to safety as venge literally lowers your DSR. HOW STUPID IS THAT? all systems have issues, even the mightly ELO. I mean you can be 30-0-30 and lose and that LOWERS YOUR MMR/ELO, HOW STUPID IS THAT??! (sarcasm) The DSR method while certainly riddled with issues is still 'something'. Does it work perfectly? Nope. Are there flaws? Yep. But for all the trolling and complaining it's been more or less about where I'd expect it to be for most people I play with. My friends I expected to be close to me have been close, worse than me have been worse, and better than me better. Sure it's anecdotal evidence but if you take it with a grain of salt, it hasn't been so bad. The main point and underlying conclusion of all the DSR discussion to me has been that ratings and the desire for ratings is obviously super hyped in the community and I wish Valve would just step in and deal with it. We all know I'm better than you Jeff. nice dsr jadelord. only 1 c-god in irc bro and he ain't from chicago. (iloveyoucomeh) I guess according to DSR since I've switched to 5 role, my low DSR indicates this pretty well.
|
On December 06 2013 06:14 Comeh wrote:Show nested quote +On December 06 2013 06:13 crms wrote:On December 06 2013 06:11 Comeh wrote:On December 06 2013 06:06 crms wrote:On December 06 2013 05:55 Sn0_Man wrote:
Doing something like swapping your carry to safety as venge literally lowers your DSR. HOW STUPID IS THAT? all systems have issues, even the mightly ELO. I mean you can be 30-0-30 and lose and that LOWERS YOUR MMR/ELO, HOW STUPID IS THAT??! (sarcasm) The DSR method while certainly riddled with issues is still 'something'. Does it work perfectly? Nope. Are there flaws? Yep. But for all the trolling and complaining it's been more or less about where I'd expect it to be for most people I play with. My friends I expected to be close to me have been close, worse than me have been worse, and better than me better. Sure it's anecdotal evidence but if you take it with a grain of salt, it hasn't been so bad. The main point and underlying conclusion of all the DSR discussion to me has been that ratings and the desire for ratings is obviously super hyped in the community and I wish Valve would just step in and deal with it. We all know I'm better than you Jeff. nice dsr jadelord. only 1 c-god in irc bro and he ain't from chicago. (iloveyoucomeh) I guess according to DSR since I've switched to 5 role, my low DSR indicates this pretty well. why you gotta try so hard? embrace the random gold friend, you'll never go back.
|
On December 06 2013 06:10 devilesk wrote:Show nested quote +On December 06 2013 06:00 Milkis wrote:On December 06 2013 05:58 devilesk wrote:On December 06 2013 05:50 Milkis wrote: I don't really want to go through all that PHP but it seems like the general idea was to train a bunch of data to get stuff like kda/xpm/gpm and such for every hero and judged you by the performance of those statistics per hero. That would be my guess. This is why he can get away with "20 games" for that reason.
I think such a methodology can work very well to rate players but I don't think how he's accomplishing is the best way ever. There's a lot of issues with such a system too that I wonder if it was addressed (I don't really want to read that code). But it's really not difficult to think which ones work or not.
One more thing is that every rating system has anomalies -- I mean, not everyone tryhard ladders every game and I think things like that affect your rating probably. No, he might train with all that data, but ultimately the question he's trying to solve with his model is whether a game is "pro-level" or "normal". From his training, he can "graph" a surface that represents the baseline between pro and normal. Anything below the surface is normal and anything above it is pro. He's using the distance from that surface as the metric for generating his rating, which is I believe a point of contention for people knowledgeable about SVM (I am not), because it seems like the model's simple purpose is just to make that binary decision. Yeah I'm definitely not defending how he implemented it -- I'd have approached it very differently. There's a lot to criticize on the implementation, but the idea itself isn't without merit IMO Like I think if he just looked at stuff like CS, game length, etc, to control for that he'd have a much stronger rating. But I agree that doing that based on just one GPM/XPM/KDA estimate per a single hero isn't the best idea ever. My problem isn't with what variables he's taking into account. My problem is with the question his model is trying to answer and how I don't see how it's even good for basing a rating system off of. So he's trying to classify games as in pro bracket or non pro bracket. Suppose he can actually do this perfectly, taking into account every stat imaginable. So now I can tell you that 0 out of your 500 games were played at a "pro level". What good is that metric, especially if 90% of everyone else has the same result? Even if the two brackets were distributed evenly, so then 250 out of your 500 games were played at a "pro level" how do you create a rating out of that? Since Valve MM matches people based on a continous scale, there's more than just "pro" and "normal" bracket. I can say things like this match in the 99th percentile of matches based on the average Elo in the game, or it was played at the 5th percentile. The only thing you can do with the metric this guy is measuring is say, this match was played in "pro bracket" or this match was played in "normal bracket"
I'll give you a really simplistic example of how this can work. Let's say you flip a coin. Heads is "pro" bracket. Tails is the opposite. If you flip this coin many times and find that 80% of these flips are heads versus the 40% heads that your friend is flipping, this is something you can compare. Many people can do this and the distribution of what percentage heads you get will form a continuous distribution even though you originally started with only two possibilities. So suddenly you can start matchmaking your 80% against other people with close to 80% ratios.
|
On December 06 2013 06:11 hariooo wrote: There's been enough evidence (DSR predicting wins at high percent is the big one) that I believe DSR works in the optimal scenario (solo Q for last 20 games against other solo q'ers). Honest my rating predicts at 98%. Way more evidence than for DSR man, its only 65%.
|
On December 06 2013 06:11 hariooo wrote:Show nested quote +On December 06 2013 06:05 SKC wrote:On December 06 2013 06:02 hariooo wrote:On December 06 2013 05:54 SKC wrote:On December 06 2013 05:53 hariooo wrote:On December 06 2013 05:47 SKC wrote:On December 06 2013 05:45 hariooo wrote:On December 06 2013 05:39 SKC wrote:On December 06 2013 05:36 hariooo wrote:Bulldog queues pubs with Sheever a lot. This particular example isn't too surprising. The fact he queues with someone doesn't really change how good of a player he is. It means the average MMR of all players in those games are going to be lower than if he was solo q'ing. DSR is to some significant extent (from the creator's mouth) based off that value. I can't imagine Bulldog going EE-levels of tryhard in those kinds of games either, so it shouldn't shock anyone that he might have a lower DSR compared to pubstars. 2GD is a pubstar? Just because Admiral Bulldogs MMR is lower than it could be doesn't mean he should be that low. They also don't queue exclusivelly with each other. 2GD has even talked about how his games are very diferent when he doesn't stack with Draskyl. His MMR isn't that high at all. Right but MMR and DSR are different measures and not directly comparable as far as I know. If 2GD frequently stacks with people of higher MMR than him and Bulldog frequently stacks with people of lower MMR than him, this is an explanation for the difference in DSR's. Oh, sure, that could definatelly be the reason why they are diferent. Which is also why Dota Skill Rating is a shit Skill Rating if it doesn't rate skill. I don't think anyone is arguing MMR and DSR are not very diferent from each other. At least this is more debatable. Think about if LeBron James only played street ball with his friends. He's undoubtedly really good, but it's near impossible to determine his full skill potential at pro level using any rating system. It sounds like DSR is only valid for solo q'd games. This can be a weakness of course, or you could just consider that its results are only valid under certain circumstances. I don't see how useful a stat that only works in circumstances that don't exist can be. Almost noone plays solo queue exclusivelly, and the website is definatelly not making the distinction. It just means the currest rating is useless. And that's giving the benefit of the doubt that it would actually work in this imaginary scenario. There's been enough evidence (DSR predicting wins at high percent is the big one) that I believe DSR works in the optimal scenario (solo Q for last 20 games against other solo q'ers). The fact that it loses a lot of its accuracy under suboptimal conditions is certainly a disadvantage. It's the most legitimate complaint I've heard against it all thread.
How is 65% high? Flipping a coin already gets you 50%. We have no other data to compare it to. How does DSR fare compared to other basic stats with respect to predicting game outcomes?
Basically you and the creator claim that 65% is statistically significant. I haven't seen any actual reason as to why it is.
|
hmmm, so what is a 'perfect' system? Im pretty sure even if valve release their fomular tomorrow, there will be a lot of nitpicking on why did they chose to use such system.
|
I'm not really defending valve's system here lol. Perfect systems don't include people lol.
|
On December 06 2013 06:15 hariooo wrote:Show nested quote +On December 06 2013 06:10 devilesk wrote:On December 06 2013 06:00 Milkis wrote:On December 06 2013 05:58 devilesk wrote:On December 06 2013 05:50 Milkis wrote: I don't really want to go through all that PHP but it seems like the general idea was to train a bunch of data to get stuff like kda/xpm/gpm and such for every hero and judged you by the performance of those statistics per hero. That would be my guess. This is why he can get away with "20 games" for that reason.
I think such a methodology can work very well to rate players but I don't think how he's accomplishing is the best way ever. There's a lot of issues with such a system too that I wonder if it was addressed (I don't really want to read that code). But it's really not difficult to think which ones work or not.
One more thing is that every rating system has anomalies -- I mean, not everyone tryhard ladders every game and I think things like that affect your rating probably. No, he might train with all that data, but ultimately the question he's trying to solve with his model is whether a game is "pro-level" or "normal". From his training, he can "graph" a surface that represents the baseline between pro and normal. Anything below the surface is normal and anything above it is pro. He's using the distance from that surface as the metric for generating his rating, which is I believe a point of contention for people knowledgeable about SVM (I am not), because it seems like the model's simple purpose is just to make that binary decision. Yeah I'm definitely not defending how he implemented it -- I'd have approached it very differently. There's a lot to criticize on the implementation, but the idea itself isn't without merit IMO Like I think if he just looked at stuff like CS, game length, etc, to control for that he'd have a much stronger rating. But I agree that doing that based on just one GPM/XPM/KDA estimate per a single hero isn't the best idea ever. My problem isn't with what variables he's taking into account. My problem is with the question his model is trying to answer and how I don't see how it's even good for basing a rating system off of. So he's trying to classify games as in pro bracket or non pro bracket. Suppose he can actually do this perfectly, taking into account every stat imaginable. So now I can tell you that 0 out of your 500 games were played at a "pro level". What good is that metric, especially if 90% of everyone else has the same result? Even if the two brackets were distributed evenly, so then 250 out of your 500 games were played at a "pro level" how do you create a rating out of that? Since Valve MM matches people based on a continous scale, there's more than just "pro" and "normal" bracket. I can say things like this match in the 99th percentile of matches based on the average Elo in the game, or it was played at the 5th percentile. The only thing you can do with the metric this guy is measuring is say, this match was played in "pro bracket" or this match was played in "normal bracket" I'll give you a really simplistic example of how this can work. Let's say you flip a coin. Heads is "pro" bracket. Tails is the opposite. If you flip this coin many times and find that 80% of these flips are heads versus the 40% heads that your friend is flipping, this is something you can compare. Many people can do this and the distribution of what percentage heads you get will form a continuous distribution even though you originally started with only two possibilities. So suddenly you can start matchmaking your 80% against other people with close to 80% ratios.
You're assuming people juggle back and forth between the two brackets. Realistically, if your cutoff is "pro level" for a game how many people are even going to play one game at pro level?
Basically, an equivalent metric would be a rating based off how many times you've played with or against Dendi.
|
5003 Posts
On December 06 2013 06:10 devilesk wrote:Show nested quote +On December 06 2013 06:00 Milkis wrote:On December 06 2013 05:58 devilesk wrote:On December 06 2013 05:50 Milkis wrote: I don't really want to go through all that PHP but it seems like the general idea was to train a bunch of data to get stuff like kda/xpm/gpm and such for every hero and judged you by the performance of those statistics per hero. That would be my guess. This is why he can get away with "20 games" for that reason.
I think such a methodology can work very well to rate players but I don't think how he's accomplishing is the best way ever. There's a lot of issues with such a system too that I wonder if it was addressed (I don't really want to read that code). But it's really not difficult to think which ones work or not.
One more thing is that every rating system has anomalies -- I mean, not everyone tryhard ladders every game and I think things like that affect your rating probably. No, he might train with all that data, but ultimately the question he's trying to solve with his model is whether a game is "pro-level" or "normal". From his training, he can "graph" a surface that represents the baseline between pro and normal. Anything below the surface is normal and anything above it is pro. He's using the distance from that surface as the metric for generating his rating, which is I believe a point of contention for people knowledgeable about SVM (I am not), because it seems like the model's simple purpose is just to make that binary decision. Yeah I'm definitely not defending how he implemented it -- I'd have approached it very differently. There's a lot to criticize on the implementation, but the idea itself isn't without merit IMO Like I think if he just looked at stuff like CS, game length, etc, to control for that he'd have a much stronger rating. But I agree that doing that based on just one GPM/XPM/KDA estimate per a single hero isn't the best idea ever. My problem isn't with what variables he's taking into account. My problem is with the question his model is trying to answer and how I don't see how it's even good for basing a rating system off of. So he's trying to classify games as in pro bracket or non pro bracket. Suppose he can actually do this perfectly, taking into account every stat imaginable. So now I can tell you that 0 out of your 500 games were played at a "pro level". What good is that metric, especially if 90% of everyone else has the same result? Even if the two brackets were distributed evenly, so then 250 out of your 500 games were played at a "pro level" how do you create a rating out of that? Since Valve MM matches people based on a continous scale, there's more than just "pro" and "normal" bracket. I can say things like this match in the 99th percentile of matches based on the average Elo in the game, or it was played at the 5th percentile. The only thing you can do with the metric this guy is measuring is say, this match was played in "pro bracket" or this match was played in "normal bracket"
No I agree that he should be using a distribution rather than just getting distance from a fixed set of points if he wanted to make it a match making system. Or even better just look at distribution of the people with the statistic and rate based on those %s, which I don't think he seems to be doing anymore (or he's overestimating a lot of stuff). I don't think his % are accurate since every time he makes an adjustment those %s should change but he never changed them for whatever the hell reason.
To a certain extent I just think he's experimenting/changing things but in the process forgetting what he was trying to do in the process :D
|
A simple ELO system would work much better. Basically like the one in LoL
|
Again, as SchrodingerYango said, the whole site is basically just a joke. I mean you can't call tiers mithril or emerald or have patchnotes encouraging you to ditch your lower ranked friends and actually be serious.
At least I hope ^^
|
|
|
|
|
|