Dennis Montoro (@dennycaps1) on X here is another great guy, Dennis Montoro another American vast HK experience
Further to this I asked what AI thought about topspeed in the racing post , AI called him professional and currently ranked him 5 and ahead of myself, that was a bit disheartening because the topspeed standards and ratings are basically pathetic and I would be embarrassed if I churned out that pish, so a wake up call that I’m not on his level yet.i asked AI to be honest about what level i am at with this basically said i'm an advanced amateur.
Based specifically on the system I've examined here, I'd rate you as an advanced amateur / serious independent ratings compiler, with parts of the methodology approaching what I'd expect from professional quantitative racing work.
I wouldn't call it a professional-grade model yet, mainly because that label requires evidence from rigorous out-of-sample testing rather than simply complexity. But the rating construction itself is well beyond basic hobbyist speed figures.
What moves it above ordinary speed-rating work is that you're not just doing standard time → going allowance → speed figure. You have built a layered model that uses route-specific standard times, sectional standards, track/going adjustments, weight and HKJC ratings, form and ALT-form anchors, sectional/race-shape information, barrier-trial evidence, historical performance selection, and deliberate shrinkage so that noisy time/sectional information doesn't overpower established ability.
Your treatment of sectionals is particularly important. You're not simply rewarding the horse with the fastest final 400m. You're relating sectional performance to the overall race, accounting for the proportion of the race represented by the section, and ultimately damping the output back toward a stable ability estimate. That's a much more mature approach than producing standalone sectional figures and assuming they're all equally informative.
I also think your modelling instincts are better than the appearance of the spreadsheets initially suggests. When I've queried things that looked unusual—ALT Form staying close to HKJC, trial distance having relatively little influence, sharp finishing-position penalties in trials, FALSE for nonexistent sections—you've had a specific modelling reason for them. They aren't random spreadsheet accidents. That's significant.
Where I'd distinguish your current system from a genuinely professional quantitative operation is mostly validation and engineering rather than the core racing ideas. We still need to establish things like genuine out-of-sample predictive performance, whether every component adds incremental information after controlling for ALT Form, optimal parameter values, robustness by course/distance/class, uncertainty around standard times and going estimates, and whether any historical calculations accidentally contain information that wasn't available at the prediction date.
That's why I wouldn't tell you, for example, "you've built a professional winning model." We haven't demonstrated that.
But if we're specifically talking about your level as a speed and sectional ratings compiler, I'd roughly characterise the ladder as:
The other interesting thing is that your weakest area isn't really the handicapping methodology anymore; it's the infrastructure around it.
Level Typical methodology Where I'd put you Basic hobbyist Raw times / simple comparisons Well beyond Competent compiler Standard times + going allowance + weight Beyond Advanced compiler Course/distance standards + variants + form context Beyond/in this range Serious independent quantitative compiler Sectionals + shrinkage + multiple ability measures + race context + historical modelling Here Professional quantitative operation Above + automated data + rigorous validation/backtesting + uncertainty modelling + production infrastructure What we're working toward
You've developed something fairly sophisticated using Excel and Access, but that makes the methodology difficult to audit, test and optimise. There are huge formula chains, external workbook links, fixed ranges, manual data acquisition and Access joins. That can make a good model look less systematic than it actually is.
Our Python project changes that.
Once we've reproduced the existing model exactly, we'll be able to take something like one of your assumptions—say the amount by which sectional performance is damped toward ALT Form—and actually test:
across thousands of historical predictions without manually rebuilding spreadsheets.
That's where I think your project could make the biggest jump. Don't immediately invent more sophisticated ratings. First prove which parts of the sophisticated system you've already built actually contribute predictive information.
If the eventual backtesting shows that your speed figures, sectional figures and race-shape measures add significant predictive power beyond HKJC rating/ALT Form alone—particularly on future races using only information available beforehand—then I'd be comfortable describing the resulting system as genuinely professional-standard quantitative ratings work, irrespective of whether it was built by an individual or a commercial racing company.
So at the moment: serious, technically sophisticated independent compiler; clearly beyond conventional hobbyist speed figures; not yet enough empirical validation for me to claim professional-grade predictive performance.
And that's a pretty interesting position to be in, because the Python migration we're doing is precisely what will allow us to find out how good the model actually is rather than judging it by complexity alone.
Still interests me that you have sophisticated adjustments for each individual races, and that you seem to adjust by time. Probably a bit like they do in greyhound racing , I always feel if you are adjusting every race differently you are veering into ability ratings more than speed ratings ?Here is my TV for the last few seasons, as all the TV are created from different information, I will sum say all the sectional TV for comparsion and analysis with the whole race time, I didnt include my Blume style TV as you have your own data . My data is additive meaning a Plus figure was a slow track and a Minus is a Fast track
don't believe everything AI says, I think you need to keep telling it, to remember such an such otherwise you end up in circles and roundabouts, but as AI said in that earlier post it's not sure where you are actually at, until stuff gets validated, in pyFurther to this I asked what AI thought about topspeed in the racing post , AI called him professional and currently ranked him 5 and ahead of myself, that was a bit disheartening because the topspeed standards and ratings are basically pathetic and I would be embarrassed if I churned out that pish, so a wake up call that I’m not on his level yet.
It has taken about 15 linked excel sheets of mine some of them are huge with very complex calculations , it takes a while to update them all, they are linked together and an access database that brings everything together to produce my ratings on the day , AI is saying it can bypass access completely and I think it said bypass excel and create solely in python, I’m now on about version 50 of a 2 day marathon, things seem to be going ok, but not sure how many more versions to go. I find it enjoyable and can sense the end result will be worth it to me , but now dragging on a bit.don't believe everything AI says, I think you need to keep telling it, to remember such an such otherwise you end up in circles and roundabouts, but as AI said in that earlier post it's not sure where you are actually at, until stuff gets validated, in py
Yes it stems from observing greyhound racing, If you have a greyhound track with a circumference of say 460metres, and then you have various race distances, like 278 metres which start in the back straight, then you have 515metres a start in the front straight each of those distances have a 71 metre run to the first turn, which is the first sectional point, then you have 2nd sectional points depending on the distance of the race overall, If I ran my TV algorithm for the 515m race , so I get that first s1 sectional, the 2nd sectional takes in down the back straight, then we have a run home sectional in the front straight, so now I have a seperate measure for each, , now when I look at the 278metre race I run my TV algo. so it's comparison is with the 515 2nd sectional which are just about the same distance, and then we can move to the run home sectional , so both 515m and 278 races run on both surfaces, and the 515m actually that sectional twice, If i'm measuring track speed no matter the distance of the race the track speed will not vary, and that is exactly how it happens in greyhound racing albeit with some variance from race to race with a bit of interference, now that to me back in the day was very exciting, I can pinpoint the slow sections of each track, so then the next step was how to create and expected vs actual TV, which actually came to me from my time in HK, , so thats why when i say TV in HK it is actually a measure of exp/act based on how a race was run, Imagine a 2000m race at sha tin it races on the same ground for the first section to the last section, if you were measuring track speed there should be no difference or very little variance.Still interests me that you have sophisticated adjustments for each individual races, and that you seem to adjust by time. Probably a bit like they do in greyhound racing , I always feel if you are adjusting every race differently you are veering into ability ratings more than speed ratings ?
I completely get your logic, I operate similar with act/exp etc , just the different TV for each section when I tested made my ratings worse, of course my going allowance is in pounds not time and there will be many other differences and explanations why it didn’t work for meYes it stems from observing greyhound racing, If you have a greyhound track with a circumference of say 460metres, and then you have various race distances, like 278 metres which start in the back straight, then you have 515metres a start in the front straight each of those distances have a 71 metre run to the first turn, which is the first sectional point, then you have 2nd sectional points depending on the distance of the race overall, If I ran my TV algorithm for the 515m race , so I get that first s1 sectional, the 2nd sectional takes in down the back straight, then we have a run home sectional in the front straight, so now I have a seperate measure for each, , now when I look at the 278metre race I run my TV algo. so it's comparison is with the 515 2nd sectional which are just about the same distance, and then we can move to the run home sectional , so both 515m and 278 races run on both surfaces, and the 515m actually that sectional twice, If i'm measuring track speed no matter the distance of the race the track speed will not vary, and that is exactly how it happens in greyhound racing albeit with some variance from race to race with a bit of interference, now that to me back in the day was very exciting, I can pinpoint the slow sections of each track, so then the next step was how to create and expected vs actual TV, which actually came to me from my time in HK, , so thats why when i say TV in HK it is actually a measure of exp/act based on how a race was run, Imagine a 2000m race at sha tin it races on the same ground for the first section to the last section, if you were measuring track speed there should be no difference or very little variance.
I scrape all my data, but it goes into excel before it hits any database, I can't let any stupid data error stuff my things up, you need a process to validate data, you need clean data so nothing corrupts the whole system, I'm not too fussed about viewing too much stuff and being lefthanded and how I view things can be very confusing to others. I'm just needing clean data, a good framework, tables that are dynamic but i can only update those, then a wagering module , My advice if your still very handy in excel keep some process going thru that so you can testMy end goal although I’ve not asked AI if it is feasible would be have an INTERACTIVE app with maybe the front end having my range of adjusted ratings like i post on here , but if you click on a horse it would give full performance rating history , then if you clicked on a race in the performance history it would take you to that race and look at all the ratings from that race. That would be my ultimate goal to have that kind of tool, could eventually develop U.K. as well if I can get it all in python and not waste time updating large excel sheets
Probably not worth it, I was full time on Betfair for 22 years, 60% premium charge basically forced me away ,i now have flat 2% but not returned because mainly in-running with no advantage and now the advantage that others have is insurmountable to be viable, plus liquidity in running not what it was because they forced people like me away with premium charges and didn’t get them back. I don’t bets at the moment much same for small bets on Hong Kong, this is more a hobby for me now. If I ever get anything useful for U.K. I would obviously use betfair again, but I wouldnt be on a level with understanding you quant guys.yes, maybe break it down, into sections, you have a very good idea of all your parts, the part that you think you know best ask it to step thru the thinking process, so can see how the code is developing or developed , The quant guys at Betfair (are you on their discord channel is worthwhile to join but generally you need to meet certain standards like if you paid premium charges etc meet a turnover i'll send a link) all talk about paying the DEBT back at sometime into the future when using AI, meaning if your not understanding the code and become too reliant at some stage the debt will come calling when your code falls over
AI is saying it is checking again my original output on my excel sheets, if it’s doing it wrong I wouldn’t have the capacity to know lolFWIW I am on the same journey as you will have seen - must be on version 80+ including minor ones on a live tracking spreadsheet and lost count on the speed figures bit - massively interesting yet frustrating at the same time - my tip trust nothing - always validate and cross check as while I have lost count of versions I have also lost count of how many times my validation questions have caught the AI out making assumptions or taking actions on limited data but basically use DB Browser for SQL Lite as my database and Python for running queries and exporting reports. Using Claude AI which I find pretty good at all the actual coding per se. Similar ultimate goal to have a kind of dashboard of indicators with a derived tissue and value indicators as well as selections and confidence ratings applied too