• Hi Guest Just in case you were not aware I wanted to highlight that you can now get a free 7 day trial of Horseracebase here.
    We have a lot of members who are existing users of Horseracebase so help is always available if needed, as well as dedicated section of the fourm here.
    Best Wishes
    AR

AI Strategy

I have used them on one of the other sheets I was testing.

Copy and pasted the Timeform comments into a cell next to the selection (s), then the look up table gives +/- to each positive/negative that appears. Can also use for race comments on each run if rating each form-line.
What kind of algorithm are you using for semantic analysis? A lookup table only gives you the values entered into the system, so what happens if the analyzed text is unique? What positive/negative values do you get in such a case?
 
What kind of algorithm are you using for semantic analysis? A lookup table only gives you the values entered into the system, so what happens if the analyzed text is unique? What positive/negative values do you get in such a case?
Your original query was

StefanBelo said:
@
DuckandDive
DuckandDive Do you use semantic analysis of race descriptions in your model?

He gave you a perfectly good answer in a public part of the forum
Most of the "good stuff" on here is available in other less pubic parts of the forum - viewing privileges are decided by the moderators but are based on fair use and a member MUST have made a positive contribution to the forum
He also explained that he only works within Excel so any "algorithms" would be "proprietary"
Why don't you post some examples of your own semantic / sentiment analysis towards "race descriptions" and explain the concept further ??

This type of analysis is used in areas such as website reviews, customer surveys and consists of analysing mainly "text" based data
In horse racing most "text" based data is standardised - an example being public "in-race" comments - race readers over the years have conformed to a certain standardisation of the process which has resulted in these type of in-race comments becoming "normalised"
They are still only one person's opinion of what happened in a race
Punters though mostly have an excellent idea of what is "positive" and what is "negative" - yes you could apply a classifying algorithm such as "Naive Bayes" that outputs a probabilistic score (usually a p-value between 0-1) to each comment - i would imagine that would require plenty of back data to be applied as a training process before the scores showed decent "separation" between the predicated classes - you could apply a 3 class model or a basic 2 -class model that spans the probabilities of "Positive" to "Negative" through a process called "vectorisation"

Bottom line though is that these "comments" are only a generalised view of how each horse ran in the race - plenty of stuff is missed by public race readers - if there is any real "edge" to be found in this area IME it requires the very hard work of watching the race replays yourself - your own personal comments and observations can be more expansive and thorough - these can then be applied to personal performance ratings and sectional and final time data -Would there be any worth in setting up such a venture ?? - YES, extremely so IME !!!! - i would also say though that unless "specialisation" (for eg - 2yo's only) is applied this would be beyond the scope of the VAST majority of punters.
 
What kind of algorithm are you using for semantic analysis? A lookup table only gives you the values entered into the system, so what happens if the analyzed text is unique? What positive/negative values do you get in such a case?
Here i've used a "naive bayes" classifier on yesterdays in-race comments based on a 2-step model that spits out a probability based on vectorisation that spans the scale of 0 to 1.00 where 0 = MOST NEGATIVE and 1.00 = MOST POSITIVE
Note - NO AI was used here, i'm a coder and well versed in python and R
Still don't see any REAL edge here though - these comments are very public - every punter in the land reads them

Here's my code

import pandas as pd
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.naive_bayes import MultinomialNB

file_path = "09.10.2025_in_race_comments.csv"
df = pd.read_csv(file_path)

comments = df.iloc[1:423, 7].dropna().astype(str).reset_index(drop=True)

train_texts =

["Positive"
"kept on well", "ran on strongly", "won easily", "stayed on", "led throughout",
"quickened clear", "stayed on strongly", "ran on gamely", "made all, stayed on well",

"Negative"
"no impression", "weakened quickly", "outpaced", "never a factor", "well beaten",
"headed final furlong", "kept on one pace", "soon weakened", "driven and weakened",
"ridden and no extra", "weakened final furlong"]

train_labels = [
"Positive", "Positive", "Positive", "Positive", "Positive",
"Positive", "Positive", "Positive", "Positive",
"Negative", "Negative", "Negative", "Negative", "Negative",
"Negative", "Negative", "Negative", "Negative", "Negative", "Negative"
]
vectorizer = CountVectorizer(ngram_range=(1, 2))
X_train = vectorizer.fit_transform(train_texts)

model = MultinomialNB()
model.fit(X_train, train_labels)

X_test = vectorizer.transform(comments)
probs = model.predict_proba(X_test)

pred_labels = model.predict(X_test)
positive_scores = probs[:, list(model.classes_).index("Positive")]

results = pd.DataFrame({
"In Race Comment": comments,
"Predicted Sentiment": pred_labels,
"Positive_Score": positive_scores.round(3)
})

output_path = "09.10.2025_in_race_comments_with_sentiment.csv"
results.to_csv(output_path, index=False)

print("Sentiment analysis complete. Saved to:", output_path)

print(results.head(10))

Sample of most "positive" scores

In Race Comment_Predicted Sentiment_Score_
made all, came clear final 2f, stayed on strongly, unchallengedPositive1.000
made all, ridden clear from over 1f out, ran on well, readilyPositive0.999
held up towards rear, not fluent 4th, 4th halfway, 3rd 3 out, pushed along on inner to dispute 2 out, soon led, ridden after last, stayed on wellPositive0.998
held up in touch in rear, 5th 4 out, headway on outer to dispute approaching 2 out, pushed along and led before last, ridden after last, stayed on wellPositive0.997
raced in third, closed on leaders 2f out, ran on well and led inside final 110 yards, won going awayPositive0.996
raced centre, made all, increased tempo 3f out, clear when edged left approaching final furlong, ran on strongly, unchallengedPositive0.995
towards rear, 7th halfway, 6th 3 1/2f out, took closer order approaching straight, 3rd 2f out, led narrowly over 1f out, ran on strongly final 100 yards, comfortablyPositive0.995
disputed, a little keen, led 2nd and made rest, 6 lengths lead halfway and further clear 4 out, well clear when not fluent last, kept on well, very easilyPositive0.991
held up in touch, jumped right 1st, 5th halfway, short of room on landing and checked 7th, pushed along in 4th when not fluent 2 out, went 2nd and challenged after last, stayed on well, just failedPositive0.991
made all, shaken up 2f out, ridden over 1f out, kept on wellPositive0.987
shied left at start, tracked leader in 2nd, led or disputed from 2nd, slight mistake 3rd, pushed along and led before 2 out, ridden and drew clear run-in, stayed on wellPositive0.987
tracked leaders, bumped 1st and not fluent 2nd, 4th halfway, 2nd after 3 out, pushed along and led 2 out, 1.5 lengths lead when steadied last, ridden and strongly pressed run-in, stayed on well, all outPositive0.974
rear of mid-division, 7th halfway, effort on outer 1 1/2f out, 6th final 100 yards, ran on closing stages to 3rd close home, nearest at finishPositive0.973
mid-division, 6th halfway, 5th 7f out, stayed on in 4th inside final furlong, ran on in 2nd closing stages to lead final stridePositive0.963
rear of mid-division, 8th halfway, took closer order over 4f out, headway on outer under 3f out, led over 2f out, asserted 1 1/2f out, clear inside final furlong, kept on well, easilyPositive0.955
held up towards rear, 6th when not fluent 3 out, pushed along in 4th before 2 out, ridden in 3rd before last, led run-in, kept on well, cosilyPositive0.944
chased leaders, switched left and went 2nd 3f out, led over 1f out, clear inside final furlong, stayed on wellPositive0.940
dwelt slightly, soon prominent, 4th halfway, took closer order under 2f out, stayed on to lead 1f out, driven out closing stages, reduced advantage close homePositive0.933



Sample of most "negative" scores

taken down early wearing hood, took keen hold, close up, shaken up over 2f out, ridden and edged left over 1f out, no extra and weakened gradually inside final furlongNegative0.000
chased leaders, ridden over 1f out, one pace and no impression final furlongNegative0.000
soon raced in second, outpaced and lost position home turn, weakened quickly final furlong, lost action and eased towards finish, dropped to last final stridesNegative0.000
led, headed after first furlong, prominent, ridden over 2f out, outpaced and detached with one other rival approaching final furlong, weakened quickly final furlongNegative0.000
hampered start, off the pace in last trio, some headway near side of group over 1f out, 5th and no impression final furlongNegative0.000
chased leaders near side of group, ridden and no impression when lost 3rd inside final furlong, weakened soon afterNegative0.000
raced far side of group, off the pace towards rear, pushed along and struggling halfway, never a factor, tailed off and eased inside final furlongNegative0.000
slowly into stride, held up in last pair, headway near side of group 2f out, went 2nd 1f out, ridden and challenging briefly inside final furlong, no extra final 110 yardsNegative0.000
close 3rd, went 2nd over 2f out, led 2f out, ridden and headed over 1f out, lost 2nd 1f out, no extra in 3rd inside final furlong, eased towards finishNegative0.000
in rear, 10th halfway, under pressure in 9th 1 1/2f out, no impression, weakened final furlongNegative0.000
tracked leaders, 4th 4 out, 5th when slight mistake next, ridden and headway to dispute before 2 out, headed 2 out, disputed again last, headed and no impression run-in, kept on one pace, 3rd close homeNegative0.000
led, joined 2nd and led or disputed after, not fluent 4 out, pushed along and headed before 2 out, ridden and no impression before last, kept on one paceNegative0.000
tracked leader in 2nd, close up at times, pushed along briefly after 4 out, ridden and no impression before 2 out where dropped to 5th, no extra, kept on one paceNegative0.000
in touch towards rear, 6th halfway, slight mistake and pushed along 8th, mistake and dropped to rear 3 out, ridden and no impression before 2 out, weakenedNegative0.000



Full results attached
 

Attachments

  • 09.10.2025_in_race_comments_with_sentiment.csv
    54 KB · Views: 3
Last edited:
Here i've used a "naive bayes" classifier on yesterdays in-race comments based on a 2-step model that spits out a probability based on vectorisation that spans the scale of 0 to 1.00 where 0 = MOST NEGATIVE and 1.00 = MOST POSITIVE
Note - NO AI was used here, i'm a coder and well versed in python and R

Here's my code

import pandas as pd
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.naive_bayes import MultinomialNB

file_path = "09.10.2025_in_race_comments.csv"
df = pd.read_csv(file_path)

comments = df.iloc[1:423, 7].dropna().astype(str).reset_index(drop=True)

train_texts =

["Positive"
"kept on well", "ran on strongly", "won easily", "stayed on", "led throughout",
"quickened clear", "stayed on strongly", "ran on gamely", "made all, stayed on well",

"Negative"
"no impression", "weakened quickly", "outpaced", "never a factor", "well beaten",
"headed final furlong", "kept on one pace", "soon weakened", "driven and weakened",
"ridden and no extra", "weakened final furlong"]

train_labels = [
"Positive", "Positive", "Positive", "Positive", "Positive",
"Positive", "Positive", "Positive", "Positive",
"Negative", "Negative", "Negative", "Negative", "Negative",
"Negative", "Negative", "Negative", "Negative", "Negative", "Negative"
]
vectorizer = CountVectorizer(ngram_range=(1, 2))
X_train = vectorizer.fit_transform(train_texts)

model = MultinomialNB()
model.fit(X_train, train_labels)

X_test = vectorizer.transform(comments)
probs = model.predict_proba(X_test)

pred_labels = model.predict(X_test)
positive_scores = probs[:, list(model.classes_).index("Positive")]

results = pd.DataFrame({
"In Race Comment": comments,
"Predicted Sentiment": pred_labels,
"Positive_Score": positive_scores.round(3)
})

output_path = "09.10.2025_in_race_comments_with_sentiment.csv"
results.to_csv(output_path, index=False)

print("Sentiment analysis complete. Saved to:", output_path)

print(results.head(10))

Sample of most "positive" scores

In Race Comment_Predicted Sentiment_Score_
made all, came clear final 2f, stayed on strongly, unchallengedPositive1.000
made all, ridden clear from over 1f out, ran on well, readilyPositive0.999
held up towards rear, not fluent 4th, 4th halfway, 3rd 3 out, pushed along on inner to dispute 2 out, soon led, ridden after last, stayed on wellPositive0.998
held up in touch in rear, 5th 4 out, headway on outer to dispute approaching 2 out, pushed along and led before last, ridden after last, stayed on wellPositive0.997
raced in third, closed on leaders 2f out, ran on well and led inside final 110 yards, won going awayPositive0.996
raced centre, made all, increased tempo 3f out, clear when edged left approaching final furlong, ran on strongly, unchallengedPositive0.995
towards rear, 7th halfway, 6th 3 1/2f out, took closer order approaching straight, 3rd 2f out, led narrowly over 1f out, ran on strongly final 100 yards, comfortablyPositive0.995
disputed, a little keen, led 2nd and made rest, 6 lengths lead halfway and further clear 4 out, well clear when not fluent last, kept on well, very easilyPositive0.991
held up in touch, jumped right 1st, 5th halfway, short of room on landing and checked 7th, pushed along in 4th when not fluent 2 out, went 2nd and challenged after last, stayed on well, just failedPositive0.991
made all, shaken up 2f out, ridden over 1f out, kept on wellPositive0.987
shied left at start, tracked leader in 2nd, led or disputed from 2nd, slight mistake 3rd, pushed along and led before 2 out, ridden and drew clear run-in, stayed on wellPositive0.987
tracked leaders, bumped 1st and not fluent 2nd, 4th halfway, 2nd after 3 out, pushed along and led 2 out, 1.5 lengths lead when steadied last, ridden and strongly pressed run-in, stayed on well, all outPositive0.974
rear of mid-division, 7th halfway, effort on outer 1 1/2f out, 6th final 100 yards, ran on closing stages to 3rd close home, nearest at finishPositive0.973
mid-division, 6th halfway, 5th 7f out, stayed on in 4th inside final furlong, ran on in 2nd closing stages to lead final stridePositive0.963
rear of mid-division, 8th halfway, took closer order over 4f out, headway on outer under 3f out, led over 2f out, asserted 1 1/2f out, clear inside final furlong, kept on well, easilyPositive0.955
held up towards rear, 6th when not fluent 3 out, pushed along in 4th before 2 out, ridden in 3rd before last, led run-in, kept on well, cosilyPositive0.944
chased leaders, switched left and went 2nd 3f out, led over 1f out, clear inside final furlong, stayed on wellPositive0.940
dwelt slightly, soon prominent, 4th halfway, took closer order under 2f out, stayed on to lead 1f out, driven out closing stages, reduced advantage close homePositive0.933



Sample of most "negative" scores

taken down early wearing hood, took keen hold, close up, shaken up over 2f out, ridden and edged left over 1f out, no extra and weakened gradually inside final furlongNegative0.000
chased leaders, ridden over 1f out, one pace and no impression final furlongNegative0.000
soon raced in second, outpaced and lost position home turn, weakened quickly final furlong, lost action and eased towards finish, dropped to last final stridesNegative0.000
led, headed after first furlong, prominent, ridden over 2f out, outpaced and detached with one other rival approaching final furlong, weakened quickly final furlongNegative0.000
hampered start, off the pace in last trio, some headway near side of group over 1f out, 5th and no impression final furlongNegative0.000
chased leaders near side of group, ridden and no impression when lost 3rd inside final furlong, weakened soon afterNegative0.000
raced far side of group, off the pace towards rear, pushed along and struggling halfway, never a factor, tailed off and eased inside final furlongNegative0.000
slowly into stride, held up in last pair, headway near side of group 2f out, went 2nd 1f out, ridden and challenging briefly inside final furlong, no extra final 110 yardsNegative0.000
close 3rd, went 2nd over 2f out, led 2f out, ridden and headed over 1f out, lost 2nd 1f out, no extra in 3rd inside final furlong, eased towards finishNegative0.000
in rear, 10th halfway, under pressure in 9th 1 1/2f out, no impression, weakened final furlongNegative0.000
tracked leaders, 4th 4 out, 5th when slight mistake next, ridden and headway to dispute before 2 out, headed 2 out, disputed again last, headed and no impression run-in, kept on one pace, 3rd close homeNegative0.000
led, joined 2nd and led or disputed after, not fluent 4 out, pushed along and headed before 2 out, ridden and no impression before last, kept on one paceNegative0.000
tracked leader in 2nd, close up at times, pushed along briefly after 4 out, ridden and no impression before 2 out where dropped to 5th, no extra, kept on one paceNegative0.000
in touch towards rear, 6th halfway, slight mistake and pushed along 8th, mistake and dropped to rear 3 out, ridden and no impression before 2 out, weakenedNegative0.000



Full results attached
Outstanding stuff 👏👏
 
Outstanding stuff 👏👏
Still don't see any REAL edge though - this type of analysis is used in marketing and website development
Here - very PUBLIC winners get high scores
&
very PUBLIC last place finishers get low scores

in case anybody doesn't get this ...in today's markets very public "good" last time out winners are vastly overbet next time as a group

race1x.jpg
Have attached the "sentiment" scores to the races now and are in the file below
 

Attachments

  • Results - race sentiment analysis 9.10.2025.xlsx
    35.9 KB · Views: 3
Last edited:
What kind of algorithm are you using for semantic analysis? A lookup table only gives you the values entered into the system, so what happens if the analyzed text is unique? What positive/negative values do you get in such a case?

What ARAZI91 ARAZI91 said...in particular,

your own personal comments and observations can be more expansive and thorough - these can then be applied to personal performance ratings and sectional and final time data -Would there be any worth in setting up such a venture ?? - YES, extremely so IME !!!! -

Look up tables are often undervalued as a way of layering ratings, and in turn probabilties/tissues, and in turn value betting decisions.

An edge could be derived on how you score +/-. Eg: Words like "quickened" should score better than a term, "ran on well".
"Won doing handsprings" "Won with head in chest" should score very well, but you never see them written anywhere, but can become your own personal your performance metric.
 
If going down this route i would expand the model as well as the "classifiers"
Naive Bayes captures a good deal of information but Logistic Regression and Support Vector Machine (SVM) type algorithms might offer up an "alternative" and capture "information" in a different light
Also "spread" the probabilistic scores now to a 3-class model that express Positive-Neutral-Negative still on the same 0-1.00 scale

Here's those 3 algorithms in a 3-Class model on the same set of results race comments from yesterday
Again No AI was used in this - strictly my own work

My Python code

import pandas as pd
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.naive_bayes import MultinomialNB
from sklearn.linear_model import LogisticRegression
from sklearn.svm import LinearSVC
from sklearn.calibration import CalibratedClassifierCV

file_path = "09.10.2025results.csv"
df = pd.read_csv(file_path)
comments = df.iloc[1:423, 7].dropna().astype(str).reset_index(drop=True)

train_texts = [

"kept on well", "ran on strongly", "won easily", "stayed on", "led throughout",
"quickened clear", "stayed on strongly", "ran on gamely", "made all, stayed on well",
"forged clear", "won going away", "kept finding more", "driven out to win",

"kept on one pace", "ran respectably", "midfield", "never nearer", "stayed on at same pace",
"no extra final furlong", "held position", "ran evenly", "plugged on", "fair effort",
"ran creditably", "not disgraced", "finished mid-division",

"no impression", "weakened quickly", "outpaced", "never a factor", "well beaten",
"soon weakened", "driven and weakened", "ridden and no extra", "weakened final furlong",
"lost ground early", "tailed off", "always behind", "weakened inside final furlong"
]

train_labels = (
["Positive"] * 13 +
["Neutral"] * 13 +
["Negative"] * 13
)

vectorizer = CountVectorizer(ngram_range=(1, 2))
X_train = vectorizer.fit_transform(train_texts)
X_test = vectorizer.transform(comments)

nb = MultinomialNB()
nb.fit(X_train, train_labels)

logreg = LogisticRegression(max_iter=1000)
logreg.fit(X_train, train_labels)

svm = LinearSVC()
svm_cal = CalibratedClassifierCV(svm)
svm_cal.fit(X_train, train_labels)


models = {
"NaiveBayes": nb,
"LogisticRegression": logreg,
"SVM": svm_cal
}

output_frames = []

for name, model in models.items():
probs = model.predict_proba(X_test)
preds = model.predict(X_test)
class_names = model.classes_
df_out = pd.DataFrame(probs, columns=[f"{name}_{cls}" for cls in class_names])
df_out[f"{name}_Pred"] = preds
output_frames.append(df_out)

combined = pd.concat([pd.DataFrame({"In Race Comment": comments})] + output_frames, axis=1)

output_path = "09.10.2025results_classifier_comparison.csv"
combined.to_csv(output_path, index=False)

print("Saved:", output_path)
print(combined.head(5))


Sample Output to CSV at a race-level

RDateROffTimeRTrackRaceIDHoseIDFinPosFldSizeIn-Race CommentNBayesNegNBayesNeutNBayesPosNBayesPredLogRegNegLogRegNeutLogRegPosLogRegPredSVMNegSVMNeutSVMPosSVMPred
2025-10-09 00:00:0013:40:00AyrAbsolut Vodka Apprentice HandicapCisco Disco (IRE)18led early, prominent, led over 2f out, stayed on well inside final furlong, won going away0.0030.0010.996Positive0.0090.0050.987Positive0.0330.0180.948Positive
2025-10-09 00:00:0013:40:00AyrAbsolut Vodka Apprentice HandicapStar Cast28looked to stumble leaving stalls, mid-division, closed from 3f out, chased winner over 1f out, no impression well inside final furlong0.9570.0400.004Negative0.6740.2080.117Negative0.6590.2120.129Negative
2025-10-09 00:00:0013:40:00AyrAbsolut Vodka Apprentice HandicapTee Aitch Aye (IRE)38held up towards rear, headway over 2f out, went 3rd and edged left over 1f out, kept on same pace inside final furlong0.1090.8670.024Neutral0.1340.6810.185Neutral0.1520.6140.234Neutral
2025-10-09 00:00:0013:40:00AyrAbsolut Vodka Apprentice HandicapPol Roger (IRE)48chased leaders, pushed along over 2f out, no extra inside final 100 yards0.9000.0920.008Negative0.4940.3740.132Negative0.4390.4160.145Negative
2025-10-09 00:00:0013:40:00AyrAbsolut Vodka Apprentice HandicapCascade Hall (IRE)58soon led, headed over 2f out, weakened from over 1f out0.6330.0500.317Negative0.5880.1100.302Negative0.5840.0680.348Negative
2025-10-09 00:00:0013:40:00AyrAbsolut Vodka Apprentice HandicapIdyllic68held up towards rear, outpaced over 2f out, soon well beaten0.8450.0490.106Negative0.7840.1030.113Negative0.7910.0930.116Negative
2025-10-09 00:00:0013:40:00AyrAbsolut Vodka Apprentice HandicapTafsir (USA)78held up in rear, outpaced 3f out, well beaten final 2f0.8620.0660.072Negative0.7510.1360.113Negative0.7510.1300.119Negative
2025-10-09 00:00:0013:40:00AyrAbsolut Vodka Apprentice HandicapSophiesticate (IRE)88prominent on inside, pushed along over 2f out, weakened from over 1f out0.3260.1030.571Positive0.4010.1460.453Positive0.4810.0730.446Negative
2025-10-09 00:00:0014:15:00AyrAltos Tequila Restricted Novice Stakes (GBB Race)Northern Brave (IRE)14tracked leader after 1f, led over 1f out, drew clear inside final furlong, easily0.8340.0560.110Negative0.2540.1640.582Positive0.1520.0880.760Positive
2025-10-09 00:00:0014:15:00AyrAltos Tequila Restricted Novice Stakes (GBB Race)Gaelic Approach24soon led, pushed along over 2f out, headed and driven over 1f out, no impression on winner inside final furlong0.9940.0030.003Negative0.6820.0760.242Negative0.5610.0350.404Negative
2025-10-09 00:00:0014:15:00AyrAltos Tequila Restricted Novice Stakes (GBB Race)Wave Power (IRE)34held up in rear, switched right and effort over 2f out, no impression on winner in 3rd final furlong, no extra final 75 yards0.8370.1620.000Negative0.3500.6020.048Neutral0.3570.5880.055Neutral
2025-10-09 00:00:0014:15:00AyrAltos Tequila Restricted Novice Stakes (GBB Race)Naughty Boy44led early, tracked leader for 1f, pushed along 3f out, no impression over 1f out, weakened inside final furlong0.9990.0000.000Negative0.9460.0220.032Negative0.8390.0300.130Negative
2025-10-09 00:00:0014:50:00AyrBacardi HandicapAhamoment (IRE)17made all, ridden clear from over 1f out, ran on well, readily0.0000.0010.999Positive0.0220.0310.948Positive0.0400.0570.903Positive
2025-10-09 00:00:0014:50:00AyrBacardi HandicapMilitary Leader27held up towards rear, headway over 2f out, soon hung left, stayed on from 1f out, went 2nd well inside final furlong, no impression on winner0.4980.1610.340Negative0.2020.1040.694Positive0.2880.0340.678Positive
2025-10-09 00:00:0014:50:00AyrBacardi HandicapJkr Cobbler (IRE)37tracked leaders, ridden over 2f out, went 2nd over 1f out, no impression on winner, lost 2nd well inside final furlong0.9910.0050.004Negative0.7690.0700.160Negative0.6970.0370.265Negative
2025-10-09 00:00:0014:50:00AyrBacardi HandicapEdgewater Drive (IRE)47held up in rear, took keen hold, some headway over 2f out, no impression, well held 4th inside final furlong0.9780.0210.001Negative0.7400.1960.064Negative0.6490.2890.062Negative
2025-10-09 00:00:0014:50:00AyrBacardi HandicapOh So Cool (IRE)57chased winner until driven and edged left over 1f out, one pace0.4660.4300.104Negative0.3900.4790.131Neutral0.4570.4790.064Neutral
2025-10-09 00:00:0014:50:00AyrBacardi HandicapGreen Valentine (IRE)67mid-division, outpaced 3f out, weakened over 1f out0.5440.3340.121Negative0.6330.2250.142Negative0.7660.1430.090Negative
2025-10-09 00:00:0014:50:00AyrBacardi HandicapBeauty Blossom (IRE)77slowly into stride, mid-division, outpaced over 3f out, well beaten over 2f out0.5270.2370.236Negative0.6200.2030.176Negative0.6920.1700.138Negative
2025-10-09 00:00:0015:25:00AyrTennents Lager HandicapWoven111held up towards rear of mid-division, switched right and headway over 1f out, ran on inside final furlong, led close home0.1510.8000.050Neutral0.0900.5720.338Neutral0.0490.4400.511Positive
2025-10-09 00:00:0015:25:00AyrTennents Lager HandicapWobwobwob (IRE)211bumped start, in touch behind leaders, driven and effort over 1f out, challenged inside final furlong, led towards finish, headed close home0.9930.0050.002Negative0.6910.1660.143Negative0.5720.2160.212Negative
2025-10-09 00:00:0015:25:00AyrTennents Lager HandicapLord Bertie (FR)311held up towards rear of mid-division, not clear run closing over 2f out, not clear run and switched left over 1f out, led inside final furlong, headed towards finish0.4210.5440.035Neutral0.1120.4720.416Neutral0.0250.4010.574Positive
2025-10-09 00:00:0015:25:00AyrTennents Lager HandicapStation X411led, ridden over 2f out, driven and edged right over 1f out, headed inside final furlong, no extra towards finish0.9980.0020.000Negative0.7330.1450.122Negative0.5940.1460.260Negative
2025-10-09 00:00:0015:25:00AyrTennents Lager HandicapArt Design (IRE)511held up in rear, not much room over 2f out, switched left and headway over 1f out, never on terms0.1110.5440.345Neutral0.1720.4620.366Neutral0.1110.5000.388Neutral

Full file attached
 

Attachments

  • Results_race_sentiment_analysis_with_all_classifiers (1).xlsx
    70.4 KB · Views: 1
What ARAZI91 ARAZI91 said...in particular,



Look up tables are often undervalued as a way of layering ratings, and in turn probabilties/tissues, and in turn value betting decisions.

An edge could be derived on how you score +/-. Eg: Words like "quickened" should score better than a term, "ran on well".
"Won doing handsprings" "Won with head in chest" should score very well, but you never see them written anywhere, but can become your own personal your performance metric.
I don't mean simply scoring the public race comments
I actually mean physically watching the replays - known as "video work"
Incorporates elements of "bias-stall/path/pace/runstyle" analysis and also subjective type comments and observations based on "energy" and "effort" - sectional time data helps a lot here
"Upgrades" etc are then applied to private handicap figures and final time and sectional ratings

Takes up a big part of my time working with M&G although had the set up anyway when i was punting.
It's a lot of graft - monotonous work and constant - usually done in the "wee small hours" lol
 
..

Sentiment Analysis Breakdown



This analysis evaluates the language used in the Timeform reports to generate a sentiment score.



1. Art Of Diplomacy (Score: 0.95) 🥇



The commentary surrounding this horse is exceptionally positive. He is on a five-race winning streak over fences, and the descriptions are filled with superlative language.

  • Positive Phrases: "stretched his unbeaten run over fences to 5", "completed a hat-trick in the style of one capable of better still", "thriving like never before", "won easily", "drew clear".
  • Negative Phrases: The few negative remarks, such as "not travelling so powerfully" or "not always fluent," are minor and often from races where he still won convincingly.
  • Conclusion: The sentiment is overwhelmingly positive, reflecting a horse in the absolute form of his life.


2. Coco Mademoiselle (Score: 0.90) 🥈



This mare's reports are also highly positive, describing a very talented chaser. The language suggests high class and significant ability.

  • Positive Phrases: "thriving", "won easily", "cruised clear", "powerfully she went through this race", "remains sure to win races over fences".
  • Negative Phrases: Negative comments are mostly related to bad luck ("unfortunate Cheltenham exit", "stumbled and unseated") or a recent narrow victory that was "far from plain sailing" but where she still "rallied well", showing toughness.
  • Conclusion: The sentiment reflects a top-quality mare whose form is strong, with unlucky incidents being the main blemishes.


3. Double Powerful (Score: 0.88) 🥉



The comments for this horse's hurdling career are outstanding, indicating massive and ongoing improvement. He is making his chase debut.

  • Positive Phrases: "has made tremendous progress", "had a remarkable 12 months", "seemingly hasn't reached his limit yet", "continues to go from strength to strength", "easily made it 5 wins on the trot".
  • Negative Phrases: The few negatives are framed as excuses for his rare defeats, such as a "late mistake" and "less-than-ideal position" costing him a race he "ought to have" won.
  • Conclusion: While his lack of chasing experience is an unknown, the sentiment from his hurdling comments is incredibly strong, suggesting he has a huge engine and could be a major threat if he takes to fences.


4. Knight Of Allen (Score: 0.65)



The reports for Knight Of Allen are positive but more measured and less effusive than for the other contenders. He is also making his chase debut.

  • Positive Phrases: "back on song", "defied a penalty in straightforward style", "produced his best effort to date", "should make at least as good a chaser".
  • Negative Phrases: Comments like "couldn't reproduce his Haydock form" and "no extra from home turn" temper the overall sentiment.
  • Conclusion: The analysis suggests he is a capable horse but the language used in his reports lacks the conviction and excitement present for his rivals.



Race Conclusion



The sentiment analysis strongly points towards the two experienced chasers, Art Of Diplomacy and Coco Mademoiselle.

  • Art Of Diplomacy emerges as the top pick due to the flawless commentary on his recent unbeaten streak.
  • Coco Mademoiselle is a very close second and a clear danger. The pace hint suggests she may be ridden more prominently to take advantage of a slow pace, which could be a decisive tactic.
  • Double Powerful is the wildcard. Based on the comments about his rapid progression over hurdles, he has the potential to be better than his current rating and could be a serious threat if his jumping holds up on his chase debut
 
1760125468478.png

What were your bets today? Here are mine, using a completely different strategy that I started discussing today. Just one feature in the data doesn't seem to be a good indicator for a profitable strategy.
 
Can anyone confirm my data on the winning percentages of horses in different odds ranges?

Odds 1.01-1.99 (Heavy favorites): ~40-50% win rate (e.g., horses at 1.5 odds win about 45% of the time).
Odds 2.00-2.99: ~25-30% win rate.
Odds 3.00-4.99: ~15-20% win rate.
Odds 5.00-9.99: ~8-12% win rate.
Odds 10.00-19.99: ~4-7% win rate.
Odds 20.00+ (Longshots): ~2-4% win rate.
 
Can anyone confirm my data on the winning percentages of horses in different odds ranges?

Odds 1.01-1.99 (Heavy favorites): ~40-50% win rate (e.g., horses at 1.5 odds win about 45% of the time).
Odds 2.00-2.99: ~25-30% win rate.
Odds 3.00-4.99: ~15-20% win rate.
Odds 5.00-9.99: ~8-12% win rate.
Odds 10.00-19.99: ~4-7% win rate.
Odds 20.00+ (Longshots): ~2-4% win rate.
StefanBelo StefanBelo
Here are results bucketed by bands of BFSP since 2011 up until yesterday for UK & IRE Horse racing markets - ALL codes FlatTurf/AW & NH(Jumps) - it's not the win% that matters so much but how the actual win% relates to EXPECTED win%
There is a REVERSE fav-longshot bias in UK & IRE exchange horse racing markets - favourites & contenders are slightly overbet and longshots are slightly underbet.

UK & IRE HorseRacing Markets - 2011-2025 (current 10/10/2025)
Odds (BFSP)BetsWinsWin%Exp.Wins.BFSPWins.Above.ExpWins.Above.Exp(normalised)Exp.Win%Act/Exp.RatioMeanBFSP(p)MeanBFSP(odds)
1.01-1.506841513475.055143.14-9.14-6.6875.1810.9980.751811.33
1.51-2.00198851112955.9711208.70-79.70-20.0456.3680.9930.563681.77
2.01-2.50277101206143.5312233.59-172.59-31.1444.1490.9860.441492.27
2.51-3.00366011320636.0813231.03-25.03-3.4236.1490.9980.361492.77
3.01-4.00877482473928.1924875.07-136.07-7.7528.3480.9950.283483.53
4.01-5.00922182038122.1020400.01-19.01-1.0322.1220.9990.221224.52
5.01-6.00959971726117.9817459.54-198.54-10.3418.1880.9890.181885.50
6.01-7.00830871277315.3712733.7339.272.3615.3261.0030.153266.52
7.01-8.00771291019313.2210335.74-142.74-9.2513.4010.9860.134017.46
8.01-9.0073558867811.808579.6998.316.6811.6641.0110.116648.57
9.01-10.0070901739610.437414.65-18.65-1.3210.4580.9970.104589.56
10.01-1210704996909.059738.20-48.20-2.259.0970.9950.0909710.99
12.01-148951169757.796827.88147.128.227.6281.0220.0762813.11
14.01-167452550216.744945.8275.185.046.6361.0150.0663615.07
16.01-186434939286.103860.9467.065.216.0001.0170.0600016.67
18.01-205713931135.452890.61222.3919.465.0591.0770.0505919.77
20.01-224514521764.822257.25-81.25-9.005.0000.9640.0500020.00
22.01-244009917714.421631.36139.6417.414.0681.0860.0406824.58
24.01-263615314514.011446.124.880.674.0001.0030.0400025.00
26.01-306626524153.642377.0937.912.863.5871.0160.0358727.88
30.01-355522117033.081656.6346.374.203.0001.0280.0300033.33
35.01-404901613502.751470.20-120.20-12.262.9990.9180.0299933.34
40.01-506892715112.191378.54132.469.612.0001.0960.0200050.00
50.01-655917410411.761183.48-142.48-12.042.0000.8800.0200050.00
65.01-80388154951.28421.0473.969.531.0851.1760.0108592.19
80.01-100383153921.02383.158.851.151.0001.0230.01000100.00
100.01-10002067758280.40848.84-20.84-0.500.4110.9750.00411243.60

Data below
 

Attachments

  • Betfair prices - 2011+.xlsx
    34.7 KB · Views: 3
Last edited:
Can anyone confirm my data on the winning percentages of horses in different odds ranges?

Odds 1.01-1.99 (Heavy favorites): ~40-50% win rate (e.g., horses at 1.5 odds win about 45% of the time).
Odds 2.00-2.99: ~25-30% win rate.
Odds 3.00-4.99: ~15-20% win rate.
Odds 5.00-9.99: ~8-12% win rate.
Odds 10.00-19.99: ~4-7% win rate.
Odds 20.00+ (Longshots): ~2-4% win rate.
Your data is a bit "off" StefanBelo StefanBelo
Betfair odds are decimal not fractional
A 1.5 chance is equivalent to a bookmakers 1/2 chance ie 1/1.5 =0.66666 and 1/2 =0.66666 and win around 65% of the time
BFSPs of 1.01-2.00 win over the long term around 60.85% of the time against an expected win % of 61.18% (average odds in that bucket would be 1.63) - Expected Wins are derived by a total summation of 1/BFSP in each bucket, then average odds are derived by (1/(ExpectedWins/CountOfBets))
 
Last edited:
Back
Top