• Hi Guest Just in case you were not aware I wanted to highlight that you can now get a free 7 day trial of Horseracebase here.
    We have a lot of members who are existing users of Horseracebase so help is always available if needed, as well as dedicated section of the fourm here.
    Best Wishes
    AR

One step Two Step Models

Playing around today installing Hermes so that a local model Gemma4:e4b can access tools through Ollama and Hermes but more importantly the promise that that Hermes can learn and retain information giving promise to a genuinely local bases 'expert' in a given domain ie feed it domain information and each time you invoke it will remember its past information and future loaded information

As a test I handed it the Sung paper on comparing one step and two step (Benter style) model performance and asked it summarise it, here is the output

This excerpt discusses a comparative study of statistical models (one-step vs. two-step conditional logit models) used to predict the winning probabilities of horses in UK races, ultimately assessing the level of market efficiency in the betting market.

Here is a detailed summary of the key findings, arguments, and conclusions:

---

### 1. Comparison of Model Performance (Kelly Wagering Strategy)
The most direct test of model accuracy was a simulated betting exercise using the Kelly wagering strategy on out-of-sample races (December 1998–August 2000).

* **One-Step Procedure Model:** Achieved a return of **0.96 percent**.
* **Two-Step Procedure Model:** Achieved a significantly higher return of **17.53 percent**.

This finding strongly suggests that the two-step model captures substantially more relevant information for predicting winning probabilities than the one-step model.

### 2. Methodological Comparison (One-Step vs. Two-Step)
The paper highlights several advantages of the two-step modelling process:

* **Handling of Multicollinearity:** A crucial argument is that incorporating odds and fundamental variables simultaneously in a single step (the one-step model) likely leads to high correlations and "unstable predictions" due to multicollinearity. The two-step model separates odds-related variables from fundamental variables, reducing this issue.
* **Clarity of Variables:** The two-step model allows fundamental variables to compete for importance independently. In the one-step model, the coefficients of fundamental variables are difficult to interpret because they are influenced by how much bettors already account for that information in the odds.
* **Practical Application (Real-Time Prediction):** The two-step procedure is the only approach viable for real-time use. Step one (developing a fundamental variable model) can be done well in advance, allowing bettors to execute step two using market odds almost immediately before the race starts. The one-step procedure is too time-consuming even for modern computers.

### 3. Market Efficiency Findings
The study offers important conclusions regarding the semi-strong form efficiency of the UK betting market:

* **Market Inefficiency:** The high profit achieved using the two-step model (17.53%) suggests that the UK bookmaker betting market **is not semi-strong form efficient.** This is surprising because the variables analyzed have been public domain for decades.
* **Information Missed by the Market:** The models identified several significant variables derived from publicly available information that the market odds do not fully account for:
* **Post-Position:** The starting stall position was found to be significant, even though it was published the day before the race.
* **Average Speed Rating:** The average speed rating of a horse in its last four races plays a significant role, but this variable requires pre-analysis, implying the betting public may not utilize it.

### 4. Overall Conclusion
The final conclusion is that the accuracy of outcome probability estimates is **highly dependent on the modelling technique employed.** While both models suggested the market was "semi-strong form inefficient," the two-step model provided a much more robust and actionable assessment, demonstrating that certain types of public information are effectively being discounted by market odds.
 
Playing around today installing Hermes so that a local model Gemma4:e4b can access tools through Ollama and Hermes but more importantly the promise that that Hermes can learn and retain information giving promise to a genuinely local bases 'expert' in a given domain ie feed it domain information and each time you invoke it will remember its past information and future loaded information

As a test I handed it the Sung paper on comparing one step and two step (Benter style) model performance and asked it summarise it, here is the output

This excerpt discusses a comparative study of statistical models (one-step vs. two-step conditional logit models) used to predict the winning probabilities of horses in UK races, ultimately assessing the level of market efficiency in the betting market.

Here is a detailed summary of the key findings, arguments, and conclusions:

---

### 1. Comparison of Model Performance (Kelly Wagering Strategy)
The most direct test of model accuracy was a simulated betting exercise using the Kelly wagering strategy on out-of-sample races (December 1998–August 2000).

* **One-Step Procedure Model:** Achieved a return of **0.96 percent**.
* **Two-Step Procedure Model:** Achieved a significantly higher return of **17.53 percent**.

This finding strongly suggests that the two-step model captures substantially more relevant information for predicting winning probabilities than the one-step model.

### 2. Methodological Comparison (One-Step vs. Two-Step)
The paper highlights several advantages of the two-step modelling process:

* **Handling of Multicollinearity:** A crucial argument is that incorporating odds and fundamental variables simultaneously in a single step (the one-step model) likely leads to high correlations and "unstable predictions" due to multicollinearity. The two-step model separates odds-related variables from fundamental variables, reducing this issue.
* **Clarity of Variables:** The two-step model allows fundamental variables to compete for importance independently. In the one-step model, the coefficients of fundamental variables are difficult to interpret because they are influenced by how much bettors already account for that information in the odds.
* **Practical Application (Real-Time Prediction):** The two-step procedure is the only approach viable for real-time use. Step one (developing a fundamental variable model) can be done well in advance, allowing bettors to execute step two using market odds almost immediately before the race starts. The one-step procedure is too time-consuming even for modern computers.

### 3. Market Efficiency Findings
The study offers important conclusions regarding the semi-strong form efficiency of the UK betting market:

* **Market Inefficiency:** The high profit achieved using the two-step model (17.53%) suggests that the UK bookmaker betting market **is not semi-strong form efficient.** This is surprising because the variables analyzed have been public domain for decades.
* **Information Missed by the Market:** The models identified several significant variables derived from publicly available information that the market odds do not fully account for:
* **Post-Position:** The starting stall position was found to be significant, even though it was published the day before the race.
* **Average Speed Rating:** The average speed rating of a horse in its last four races plays a significant role, but this variable requires pre-analysis, implying the betting public may not utilize it.

### 4. Overall Conclusion
The final conclusion is that the accuracy of outcome probability estimates is **highly dependent on the modelling technique employed.** While both models suggested the market was "semi-strong form inefficient," the two-step model provided a much more robust and actionable assessment, demonstrating that certain types of public information are effectively being discounted by market odds.
Mark, the bit that interests me most is the idea of Hermes becoming a genuine local expert rather than just a document summariser.

Have you tested whether it can retain several papers or research findings over time, then later connect them correctly without being prompted with the original documents again?

For example, could you give it the Sung paper now, another paper on market efficiency next week, then some of your own racing research after that, and later ask it where the findings agree or conflict?

That seems to me the real test of whether it is actually becoming useful as a persistent domain expert rather than just doing good retrieval and summarisation.
 
Mark, the bit that interests me most is the idea of Hermes becoming a genuine local expert rather than just a document summariser.

Have you tested whether it can retain several papers or research findings over time, then later connect them correctly without being prompted with the original documents again?

For example, could you give it the Sung paper now, another paper on market efficiency next week, then some of your own racing research after that, and later ask it where the findings agree or conflict?

That seems to me the real test of whether it is actually becoming useful as a persistent domain expert rather than just doing good retrieval and summarisation.
I will let you know but first I have to hook in Osidum which will be the vehicle for storing documents but you make a very interesting point. My partner asked why you dont just ask chatGPT and I pointed out the web based LLM's will give you the conventional narrative so for example if you hold contradictory views or indeed papers it wont access them.
 
Back
Top