Historically, my teams across Europe, Asia or Central America consistently maintained large amounts of data to feed predictions and forecasts. This operational data was continuously collected across most parts of the value chain: shipment history with actual vs. promised dates, last-mile deliveries with defects upon receipt, demand by SKU, and any other metrics you can think of.
What I observed was that the answers to those questions, almost everywhere, were in spreadsheets or were based on an operator's judgment. Alternative methods also existed, but getting from a table to a proper forecast or prediction was always a painful project.
Such projects have to be budgeted, consider cross-functional resources, and require access to resources that everyone is competing for. I remember a case when my function required a specialist to run a standard ETL pipeline and train the model. This also required keeping the model alive as the data drifted. Requests like this keep data scientists busy for at least a few weeks, and my team had to wait three months to get the person available for our initiative.

My experience with probabilistic chaos and LLMs
For me, the outcomes of LLM tests were always quite disastrous. Yes, I know that any language model is powered by a probabilistic engine because it predicts the next token. It's certainly no surprise that no LLM can predict operational delays across the supply chain. What further supports my anti-LLM thesis is that the confidence any LLM highlights never correlates with how often it's actually accurate.
I've built many probabilistic agents, so I can understand their strengths and weaknesses and identify areas where they can contribute. The issues start the moment I ask for a forecast. Why? Because it's just a confident guess. In my view, the LLM's job should be just to ask for the prediction, get it, and define actions based on it.
This is how the project harness idea came to mind. The objective was shaped by the task of just keeping those two roles separated. Lately, I connected this idea to zero-shot models, which I further explore in my project.

What changed: how it helps and what accelerates the approach
First thing to admit, I'm quite amazed by zero-shot models' potential. They are pre-trained on synthetic and real tables. My current understanding, in simple words, is that they take a table they've never seen and produce a calibrated prediction without any training run. But whether those intervals are realistic is a separate validation point that the project aims to measure.
I have some hopes, especially that this design will bring me as close as possible to classical ML outcomes. But hope is one thing. I still have to validate the concept, confirm it experimentally, and report the final results.
To prepare the project, I've spent a few weeks selecting a model. Most of the candidates I've preselected include their own libraries and definitions of what a prediction or forecast is. In the middle of this process, I realised that moving from experiment to solution will take many hours of work. The main reason is that, to compare the zero-shot model with a classical baseline, the outcomes need two separate code paths. That is why I will put everything under a single user interface.
My expectations from the project
Any supply chain operator logs in, defines the columns to predict, and gets results with certain confidence intervals. No resource overkill, feature engineering, or any complex ETL required. It should be cheap and efficient (we are moving to a world where decisions become assets, not data or models themselves, right?).
Operational realities will also shape how we execute specific tasks and responsibilities. I expect a shift in how individuals approach shipment status - the agent will do all the messy chasing. At the same time, insights-powered action would stay human for quite some time. That's also why autonomy will be earned through a track record rather than as a default setup.
This will form new habits and expectations. Any solution will be expected to include built-in connections, including MCP, which I will deploy in this project as well.
I will focus mainly on a few cases/problem statements. But this could be extended further depending on how the project matures and my time availability over the next four weeks.
First is tabular prediction, which is well-positioned for operational cases such as identifying shipment delay risks. And yes, it can apply to many other areas across logistics, procurement, and manufacturing. I suspect cases like vendor scoring or OTIF issues for each shipment would fit it well, too. But I will leave some of this outside of the project, letting other peers test and conceptualise it further.
The second will be time-series forecasting. That one gets super useful for SKU-level demand patterns and inventory management. Would also be interesting (subject to my time capacity) to challenge it in a freight rate development simulation.
How I will judge the completion
At the moment of writing this update, I want to nail just four components and see how those will perform (but this might change during the course of development):
- As mentioned earlier, one interface for any task/objective irrespective of backend: zero-shot or classical
- To have a confidence interval for each prediction/forecast outcome. The main reason is that estimates without ranges are likely to overstate the numbers
- MCP connector availability due to reasons described earlier
- I expect no hidden states here, as predictions are going to be functions of the tables and overall solution configuration. Built-in backtests will help me to validate overall harness/model calibration.

I chose the project name "Auspex" because it represents, for me, the ancient meaning of someone who reads the environment and predicts what's coming.