• Hi Guest Just in case you were not aware I wanted to highlight that you can now get a free 7 day trial of Horseracebase here.
    We have a lot of members who are existing users of Horseracebase so help is always available if needed, as well as dedicated section of the fourm here.
    Best Wishes
    AR

Hong Kong Speed Figures

I always will set AI some tasks again that i know the answer. Then I set if a few tasks, it might be to make the excel formula more efficient etc, In the table I will post is just some random analysis of ST1200 what I'm looking at is over a set period of time, I want to match runners from St1200 that have had two starts at the distance and then 1 start at another distance, I'd expect AI to suggest only running analysis on three of those tables, meaning the last table the slowest runners create too much noise. I'd expect AI to be creative and say you should run Ratio's and Linear Regression and then to say run with ratios, but then use Linear Regression on those Ratios to come up with a method to adjust a runners time from one distance to another, this is when I check the outputs against what I have to what AI creates so when my tables go to py or whatever so nothing changes to what I want , so i'm fully understanding the build.
View attachment 171337
You come at the problem in a completely different way you adjust the times , from what i understand you are saying horse x has run this time and this time in ST1200 what time will horse x run in ST1400 etc. Most people myself included will adjust via the rating/ standard times, horse x will rate 100 using our speed rating calculation method in 69.226 in Sha Tin1200 and will rate 100 in 82.229 in SHA tin 1400 and bring them together that way . So it a bit like I would get the standard times to do the work. You would get the times to do the work ?
 
You come at the problem in a completely different way you adjust the times , from what i understand you are saying horse x has run this time and this time in ST1200 what time will horse x run in ST1400 etc. Most people myself included will adjust via the rating/ standard times, horse x will rate 100 using our speed rating calculation method in 69.226 in Sha Tin1200 and will rate 100 in 82.229 in SHA tin 1400 and bring them together that way . So it a bit like I would get the standard times to do the work. You would get the times to do the work ?
Yes, I would normally not use any conversion to a scaled rating like a 100, as I don't want to compress/distort the actual time differences between runners.
Using those tables I posted the true science behind it is using multiple Regression, Same as greyhounds fundamentally those runners in the top quartile will actually be specialist at their distance, so example the HV1000 cohorts might only be good at that distance but within that group are another set , that need a bit further, so using multiple regression you include the Run Home sectional when your doing your regression, so Y is final time ST1200 , and now for X axis you use HV1000 final time with F400 time , so then you start separating out the speedy squib types, but for HV1650 you will use the first say 800, with the Final time and then you get an optimal formula, but again I also use a sectional weighted TV which then is just an adjustment figure which could be added or subtracted from a Par time either LR adjusted or left raw at the same form line distance. But as you venture into a proper database you will get those opportunities to test things, I mean fundamental testing not like back fitting data, which you will get a chance to do as well
 
Yes, I would normally not use any conversion to a scaled rating like a 100, as I don't want to compress/distort the actual time differences between runners.
Using those tables I posted the true science behind it is using multiple Regression, Same as greyhounds fundamentally those runners in the top quartile will actually be specialist at their distance, so example the HV1000 cohorts might only be good at that distance but within that group are another set , that need a bit further, so using multiple regression you include the Run Home sectional when your doing your regression, so Y is final time ST1200 , and now for X axis you use HV1000 final time with F400 time , so then you start separating out the speedy squib types, but for HV1650 you will use the first say 800, with the Final time and then you get an optimal formula, but again I also use a sectional weighted TV which then is just an adjustment figure which could be added or subtracted from a Par time either LR adjusted or left raw at the same form line distance. But as you venture into a proper database you will get those opportunities to test things, I mean fundamental testing not like back fitting data, which you will get a chance to do as well
On day 3 and over 70 versions of AI building my HK ratings system in python, one piece of advice re having the option to view in excel any data scraped and added to historical data , is useful advice to me , thanks for that little nugget , makes a lot of sense, as I know clean data when I look at it.
Obviously I am going to stick with my methods at the moment, but I can see you are trying to give me sophisticated advice also on methods, when I finish building my method and AI tries to improve it with testing, I will mention some of the procedures you talk about, I guess though want to initially carry on with my methods and test them first.
 
Yeah just stick to your Plan don't deviate
It quite interesting actually, when you are talking with AI and different things pop up, I always worry about the IP of it all, I keep on asking at different times the same question but not in the same chat seeing if it will change like last night the HV1000 race so you have three sectionals for both HV1000 and HV1200 -0.196 -0.310 0.408 they are the ST1200 so how would you adjust them down for a 1000 metre race, AI compelling argument is its more race shape then distance the adjustment should reflect the race shape not the distance difference so it will always want to measure against, first measure against a par and get a zscore for each sectional , then replicate the zscore back into the 1000 race and so it would go on
 
Yeah just stick to your Plan don't deviate
It quite interesting actually, when you are talking with AI and different things pop up, I always worry about the IP of it all, I keep on asking at different times the same question but not in the same chat seeing if it will change like last night the HV1000 race so you have three sectionals for both HV1000 and HV1200 -0.196 -0.310 0.408 they are the ST1200 so how would you adjust them down for a 1000 metre race, AI compelling argument is its more race shape then distance the adjustment should reflect the race shape not the distance difference so it will always want to measure against, first measure against a par and get a zscore for each sectional , then replicate the zscore back into the 1000 race and so it would go on
Yes I find it interesting myself, thought understanding is limited, waiting for it to run a version is probably hardest part, if I was more efficient and confident I would have 1 or 2 other chats open with projects for U.K. ratings that I have concrete ideas for on the go simultaneously, but just deciding on concentrating on the one thing atm.
 
Neem thinking about the AI building tools things, my thoughts were anyone can do this if the have data they can get AI to build them a model to rate for anything they like, AI would do it the generic way with information in the public domain, the real uniqueness will only come if you have your own ideas I have solid excel sheets for creating standard times ( a very difficult task for U.K. racing , AI would probably just generically use medians or averages with the data, maybe I’m not giving it enough credit) , mine would bemore advanced I would have thought, giving AI that sheet and asking it to perform the iterations required and populating the end results could probably If I’m correct be a 1 or 2 button job, instead of spending hours with calculation heavy excel sheets. Same with creation of going allowances and speed ratings they would use methods in the public domain instead of proprietary methods. That is where if you understand racing, and the data from a racing point of view you could be adding value above what can generically be produced, that is the hope anyway
 
Back
Top