Ulric
Book a call

Eugene, Oregon · one person, whole builds

Case study

Jev Setup Score: a score that shows its own record

Type a ticker and get a 0 to 100 score for a higher close after 5, 10 or 20 sessions. Every input is on the page, and so is a walk-forward record that says plainly when the model does not beat the base rate.

ClientUlric studio product
Year2026
ScopeProduct design, Full-stack development, Machine learning, ASP.NET Core, Angular, three.js
Codegithub.com/erichers/jev-setup-score ↗
Jev Setup Score: a score that shows its own record
13model inputs, each shown on the page
504training sessions per walk-forward window
63test sessions per window
15tickers in the cache

Why I built it

A lot of market tools show a score: a gauge, a letter grade, a buy or sell meter. Very few of them show how the score was made, and almost none show how the score did in the past. I wanted to build the opposite. A score where every input is on the page, and where the model's own record sits right next to it, including when that record is not good.

This is the most direct version of the habit I call Jev decision-making. The app commits to a probability, shows what it was based on, and is scored against what happened later. The scoring is not a separate report. It is on the same page as the number.

It is an educational tool. It is not financial advice, and it is not a signal to buy, sell or hold anything.

What it does

You type a ticker and pick a horizon of 5, 10 or 20 sessions. The app returns a score from 0 to 100, which is the model's probability of a higher close at that horizon, times 100. Under it, a set of bars shows what is pulling the score up or down, each with a plain-English read, like "RSI is 60.5, leaning up" or "Volume is light, at 0.67 times the 20-day average."

Further down are a price chart with moving averages, RSI and MACD panels, and the walk-forward test: hit rate, Brier score, a calibration chart, the base rate, and an equity curve for a simple rule against buy and hold. The home page lists real past scores, the newest out-of-sample call for each name and what the close did afterward. There is a methodology page with every formula and a PDF report of each score.

How it's built

The API is ASP.NET Core 8 and the front end is Angular with standalone components and signals. EF Core stores bars and lookups in SQLite by default, or MySQL through Pomelo. ECharts draws the charts and QuestPDF writes the report. The methodology page has a three.js surface of the logistic probability over two inputs, loaded only with that view.

The model

The model is a logistic regression written in C#. It uses no paid data and no paid model. There are thirteen inputs, all from daily bars plus a weekly resample: moving-average slopes, whether the 20, 50 and 200 day averages are stacked, RSI(14), the MACD histogram and its slope, ATR as a percent of price, volume against its 20-day average, distance from the averages, and the 10-session return. Each input is standardized, so every weight is per one standard deviation, and that is what makes the bars on the page comparable.

The label is simple: is the close H sessions ahead higher than today's close? A feature row for a given day uses only bars through that day. Training is Newton's method on a penalized log likelihood, and there is one model per horizon.

The walk-forward test

The record on the page comes from a walk-forward test. The model trains on 504 sessions, about two years, and is tested on the next 63, about a quarter. Then the windows roll forward and it happens again. The rule that matters most is that a training row only counts if its answer was known before the test window starts. A 10-session label from the last days of a training window would peek into the test window, so those rows are dropped.

The xUnit suite was at 13 passing tests at the last run. Fewer tests than the other apps, but they cover the parts that would quietly make the record wrong: RSI, MACD and ATR against known series, logistic convergence, that features never look ahead, that a walk-forward row is trained only after its label is known, that all 15 cache files are ordered daily bars, and that the MySQL migrations stay on 5.7-safe SQL.

A design review round

Each build went through a design reviewer at 390, 768, 1024, 1280 and 1440 pixels wide, in light and dark, with reduced and normal motion, against a written rubric.

  • The probability surface on the methodology page waits for a Rotate press before it takes a drag, so a swipe scrolls the page on a phone. It also no longer drifts on its own after the page loads. Once it settles, the camera holds still until you press Rotate.
  • With WebGL off, or after the browser drops the 3D context, the frame shows a still of the same surface, grid and paths from the same camera. The live view comes back when the context does.
  • Raising the camera 10 degrees on wide screens made the surface easier to read, but it left 134 pixels empty on each side at 1280 and 1440. The drawing filled 73 percent of the frame, and the rubric asks for 75. The final build fits the camera to the surface, and the drawing now fills about 85 percent of the frame.

What I learned

The base rate is the number to beat, not 50%. If up closes are common in the test period, a model that always says up can look good on hit rate alone. So the page puts the base rate right next to the hit rate, adds the Brier score and a calibration chart, and the copy says so when the model does not clearly beat the base rate. I would rather the app say that plainly than leave it out.

Leakage is easy to create and hard to see. The walk-forward split looks correct at a glance even when it is not, because the leak is only a few rows at the edge of each window. That is why the rule has its own test, and why the diagram above shows the dropped rows as their own sliver.

Showing every input changed how I read the score. When I can see that most of a score comes from one input, like a tight daily range, I read the number very differently than when it is spread across many. A single gauge hides that.

TypeSafe, Jev, and where I think this goes

Jev Setup Score does not call TypeSafe's Jev model. It is a logistic regression I can read line by line. But it is built around the same idea: answer with a probability, show what it is based on, and keep score.

This section is my opinion and my prediction, so I want to label it that way.

When I say Jev decision-making, I mean a model that is asked a typed question and has to commit. It gives a probability and a call, and it is clear about what the call was based on. That answer is logged with the time and the inputs, and later it is scored against what actually happened. Jev is the name TypeSafe gave its first System One model, and it does the first half of that directly. It takes a typed question and returns typed values and probabilities, such as a choice from a list, a score on a rubric, or the probability that a statement is true, instead of a paragraph. The logging and the scoring are the half I build around it.

In the early days of ChatGPT, I felt like I could get a model to give me a number and stand behind it. That was an early capability, and in my experience it faded in the versions that followed: more hedging, more "it depends," fewer committed numbers. I do not know the full reasons for that, and I am not claiming to. It is how it felt from my side of the screen.

My view is that frontier and open model providers will build Jev and TypeSafe style typed probabilistic decision-making into the core of their models. Software needs answers it can branch on, and a typed probability I can score later is more useful to me than a paragraph I have to interpret. TypeSafe and Jev are the clearest version of that idea I have seen so far, and I think it will end up as a standard feature rather than a niche one. That is a prediction, not something I can prove today.

None of this is financial advice. Where I use Jev near markets, it is on a paper trading account with no real money, and I do not publish trading results.

What's next

  • Keeping a live log. Each lookup is already stored in the database. The next step is to score those lookups once their horizon passes, so the record grows from real use and not only from the backtest.
  • More names. The cache covers 15 tickers today.

Tech stack

  • ASP.NET Core 8 Web API
  • Angular standalone components
  • Logistic regression written in C#
  • EF Core with SQLite or MySQL
  • Yahoo Finance daily bars, then Stooq, then cached files
  • ECharts for the charts
  • QuestPDF for the report
  • three.js for the probability surface on the methodology page

Repository

The GitHub repository for this app is github.com/erichers/jev-setup-score.

Walk-forward backtest with calibration chart
Walk-forward backtest with calibration chart
Probability surface on desktop and phone
Probability surface on desktop and phone
Will the close be higher page on desktop and phone
Will the close be higher page on desktop and phone
Score readout on desktop and phone
Score readout on desktop and phone
Score page on three phones
Score page on three phones
Dark mode score page
Dark mode score page

Common questions

What does the score mean?

It is the model's probability of a higher close after 5, 10 or 20 sessions, times 100 and rounded. A 61 means the model puts that chance at about 61 percent. It is not a recommendation.

How do you know the model is not cheating?

The walk-forward test trains only on rows whose answers were known before each test window starts, and a test checks that rule. The features never use a price from after the day they describe.

Does it beat the market?

I do not make that claim. The page shows hit rate, Brier score and calibration next to the base rate, and says so when the model does not clearly beat the base rate.

← All work