A swarm of specialist agents debates probabilistic returns, correlations, and fund picks to meet your target — every claim traced to its data, every prediction scored on confidence.
Base currency INR · Indian resident investor. Pick any date — past dates auto-run a backtest against realized NAVs; today or a future date runs a forward forecast.
Run the council against your own house view instead of a generic market view. Every agent reads the compiled brief; explicit min/max allocations in it override the default caps.
The Research Council is not a single model answering a prompt. It is a 30-agent research pipeline that runs in nine phases — a macro strategist sets the world view, seven asset-class specialists argue their books in parallel while an explorer hunts wildcards and a news scanner sweeps recent events, seven red-team critics stress-test each specialist (with a revision round for anything flagged), a deterministic optimiser solves for weights against an empirical correlation matrix, seven fund-research agents build the actual scheme-level shortlist, a verifier / risk / predictions trio triangulates the portfolio from three angles, a compliance auditor gates every material factual claim, and finally a CIO synthesises the call before the reporter writes the brief. Every claim is grounded in data pulled at runtime — Yahoo prices, FRED macro series, mfapi NAV history, and primary-source scrapes of RBI, SEBI, AMFI, and AMC factsheets — and every tool call, hypothesis, and revision is logged so you can click through from any number on the page to the agent, the tool call, and the source document that produced it.
You give it a target return, a horizon, and an as-of date. If the date is in the past, the system runs a look-ahead-safe backtest and scores its own P10/P50/P90 bands against realised NAVs. If the date is today, it produces a forward forecast. Either way, you get a probabilistic view of what a portfolio can plausibly deliver — with the debate that produced it fully in the open.
One Macro Strategist forms the shared brief — India growth, inflation, RBI path, global rates, USD/INR, top risks — that every downstream agent reads first. Its output hash pins every downstream cache so a shifted view invalidates stale forecasts.
7 asset-class Specialists each produce a P10/P50/P90 with drivers, risks, and cited evidence. The Explorer hunts non-obvious cross-asset drivers and historical analogies. The News agent scans the last 30 days and tags impact by asset class.
One Critic per asset class checks each specialist for consistency with historical vol, evidence coverage, over/under-confident spreads, and missed risks. A revise verdict triggers a second round for that specialist only — accepted forecasts flow forward unchanged.
The system builds an empirical correlation matrix from historical NAVs of proxy schemes, then a deterministic mean-variance solver picks the weights that best hit your target return under realistic caps. Same forecasts, same allocation, every time.
One agent per asset class turns the specialist's shortlist into a verified fund-level book — 2+ alternatives per category, side-by-side quant comparison on returns, vol, drawdown, expense ratio, AUM, manager tenure, style, and top holdings from AMFI + AMC factsheets.
Three parallel lenses on the same portfolio: Verifier audits data integrity and numeric sanity end-to-end, Risk runs 5+ adverse scenarios grounded in historical drawdowns, Predictions forms independent per-class medians so you see where an outside view disagrees with the specialists.
Samples 15–25 of the most material numeric and factual claims across every agent, classifies each, attaches source URL + retrieval timestamp, and flags fabrications. Verdict pass / warn / fail with itemised issues.
Reads macro, news, specialists, explorer, fund research, optimizer, verifier, risk, and compliance, then produces the final allocation, rationale, and rupee-denominated action plan that sums to your initial investment.
Turns the whole debate into a clean executive summary, allocation narrative, per-asset reasoning, and consolidated key risks in plain language — with verifier and compliance caveats surfaced when material.
Every agent is specialised, prompted for its own job, and constrained to output a structured schema — so downstream agents can rely on the shape, not on hoping the previous LLM was well-behaved. Counting per-class agents individually, a single run invokes 30 LLM agents plus one deterministic optimiser: 1 macro + 7 specialists + 1 explorer + 1 news + 7 critics + 7 fund-research + 1 verifier + 1 risk + 1 predictions + 1 compliance + 1 CIO + 1 reporter.
Sets the regime.
Forms an evidence-backed view of Indian growth, inflation, RBI policy path, global rates, USD/INR, and the top global risks. Pulls FRED series and primary-source releases, argues its own regime label, and hands every downstream agent a shared brief so specialists don't reason in a vacuum. Its output is hashed into every downstream cache key.
→ Regime label, growth/inflation/rate views, INR trajectory, risks
One agent per asset class, run in parallel.
Large cap, mid cap, small cap, micro cap, international equity, gold, and Indian debt each get a dedicated agent. Each pulls its own price history, computes annualised vol and drawdowns, adjusts for the macro brief, and produces a probabilistic P10 / P50 / P90 annualised return with drivers, risks, cited evidence, and a starter fund shortlist.
→ P10 / P50 / P90 annualised return, drivers, risks, fund shortlist
Finds what the specialists miss.
Runs alongside the specialists with maximum creative freedom. Hunts non-obvious drivers, historical parallels (which decade in which country looks like India today?), cross-asset correlations (copper → cyclicals, dollar liquidity → smallcaps, real rates → gold), and second-order effects. Every hypothesis is logged with the hypothesis tool BEFORE being checked with real data — or flagged as a data gap.
→ 2–5 non-obvious hypotheses, data gaps to close
Last 30 days, in parallel with specialists.
Sweeps recent market-moving events, deduplicates, and tags each by asset-class impact. In backtest mode it is strictly capped at the cutoff date — no post-cutoff information leaks into the reasoning. Downstream agents (CIO, compliance, reporter) read this brief so nothing material from the last month is missed.
→ Ranked recent events with per-asset-class impact tags
Stress-tests every specialist.
One critic per asset class. Checks whether P10/P50/P90 are consistent with historical vol, whether claims have evidence, whether the spread is over- or under-confident, and whether an obvious current risk was missed. Votes accept or revise — and revise triggers a second round for that specialist only, so accepted forecasts flow forward unchanged.
→ Accept / revise verdict, specific critiques, suggested adjustments
Deterministic, not an LLM.
First builds an empirical correlation matrix from historical NAVs of proxy schemes per asset class, then a mean-variance solver picks the weights that best hit your target return under realistic caps and a smoothness prior. Deterministic on purpose — the same specialist forecasts always produce the same allocation, so the debate above it is what actually changes.
→ Portfolio weights per asset class, correlation matrix, expected return band
Class → actual scheme shortlist.
One agent per asset class, running in parallel after the optimiser. Takes the specialist's initial picks, adds 2+ alternatives per category, and produces a side-by-side quant comparison on trailing returns, vol, drawdown, expense ratio, AUM, manager tenure, style, and top holdings — all sourced from mfapi NAV history, AMFI, and AMC factsheets scraped at run time.
→ Verified scheme picks with quant comparison and cited sources
Data-integrity audit.
Independently audits everything before it reaches you. Checks data integrity (correct tickers, fresh macro series, currencies handled), numeric sanity (P10<P50<P90, spread plausibility, weights sum to 100), evidence coverage, and cross-consistency between the macro view and specialist calls. Can spot-check any suspicious number by re-pulling it with tools.
→ pass / warn / fail verdict, itemised issues with fixes
Adverse-scenario stress test.
Runs at least 5 adverse scenarios against the proposed portfolio — sharp INR depreciation, a Nifty drawdown regime, a global equity shock, a duration spike, a gold reversal — and grounds magnitude estimates in historical drawdowns via tools. Surfaces which sleeves carry the most tail-risk contribution.
→ Scenario matrix with drawdown estimates and mitigation notes
Independent second opinion.
Forms an independent per-asset-class median and confidence WITHOUT copying the specialists — using the same macro brief and correlation matrix as reference only. This gives you an outside view alongside the specialist debate, so you can see where the two disagree and why.
→ Independent per-class median return and confidence
Factual gate on every claim.
Samples 15–25 of the most material numeric and factual claims across every agent, classifies each (verified / partially verified / unverified / fabricated), attaches source URL + retrieval timestamp, and issues a pass / warn / fail verdict. This is the last check before the final synthesis.
→ Per-claim verification with sources, overall pass / warn / fail
Final synthesis and rupee action plan.
Reads macro, news, all specialists, the explorer, fund research, the optimiser, verifier, risk, and compliance — then produces the final call: refined allocation, rationale, key trade-offs, and a rupee-denominated action plan that sums exactly to your initial investment. Text-only (no tools) so it can't drift from the audited numbers.
→ Final allocation, rationale, per-sleeve rupee action plan
Translates the debate into a human brief.
Turns the whole council output into a clean executive summary, allocation narrative, per-asset reasoning, and consolidated key risks in plain language. If verifier or compliance flagged anything material, those caveats are surfaced under key risks — the report never quietly launders a warning.
→ Headline, executive summary, per-asset reasoning, key risks
Agents aren't boxed into a fixed template. They call tools as many times as they want, in any order — but every call is logged, every claim must be backed by evidence, and the whole trace is inspectable per run.
Ground every number in real data.
Pulls monthly price history from Yahoo Finance for any ticker and returns annualised return, vol, max drawdown, and trailing 1y / 3y / 5y. Any P10/P50/P90 that isn't anchored to this is treated as unsupported.
Fresh macro, not stale training data.
Pulls FRED series — US Fed funds, US 10y, US CPI, USD/INR, India CPI, India policy rate. Returns latest value plus 24 months of history so the agent can reason about direction, not just level.
First-hand documents, not journalist summaries.
Restricted search across a curated allow-list of PRIMARY sources — RBI, SEBI, MOSPI, PIB, Union Budget, AMFI, NSE/BSE, CCIL, Fed, BLS, BEA, ECB, BoE, BoJ, IMF, World Bank, OECD, BIS. Agents are required to pull the actual policy statement, CPI release, budget speech, or fund SID/factsheet and quote from it before making a regulatory or structural claim. Every scraped URL is logged as evidence and shown in the trace.
Go from asset class to actual fund — with fund-level P10/P50/P90.
mfapi.in gives NAV history for every Indian mutual fund (~10 years). fund_factsheet searches AMFI + the AMC's own domain (HDFC, ICICI, SBI, Nippon, Axis, PPFAS, Motilal Oswal, DSP, Mirae, Kotak, Franklin, Aditya Birla, UTI, Quant, etc.) for the primary factsheet / SID and scrapes it. Specialists compare 3-5 candidates on trailing returns, vol, drawdown, expense ratio, AUM, manager, top holdings, and style — then produce a fund-level probabilistic forecast (P10 / P50 / P90 + expected alpha vs benchmark). The backtest scorecard scores every fund pick against its OWN band, not just the class benchmark.
Reach beyond APIs when primary sources aren't enough.
Firecrawl-backed general web search and full-page scrape. Used for brokerage outlooks, sell-side reports, news events, and secondary commentary. Preferred order: primary_source_search first, then web_search only if the first-hand document doesn't cover it.
Say the wild thing out loud — then check it.
Agents record a creative hypothesis (a historical analogy, a cross-asset correlation guess, a regime-shift bet) BEFORE they check it with data. This makes creative reasoning first-class in the trace, so you can see the wildness and the falsification attempt side by side.
Auditable thinking, not black-box output.
A private note tool. Agents write out their working — what they're considering, why they're rejecting an angle, what would change their mind. Every note is logged so the whole chain of thought is inspectable, not just the final answer.
Admit what's missing.
When an agent needs data we don't have — a specific ETF, an internal Scripbox brokerage view, a sentiment index, FII/DII daily flows — it logs a request for you to plug in later, notes the caveat in its output, and continues with the best available proxy instead of pretending.
Every number on screen is traceable to a data source pulled live at run time — no cached summaries, no training-data hallucinations. The council reasons from a floor of free, primary, and API-grade feeds today, and clearly surfaces what it wishes it had via the request_data_source tool so you know exactly what a paid feed would unlock. Below is the full landscape, organised by what is wired now, what would help next, and what an institutional-grade setup would add.
The floor every run reasons from today. All fetched at run time, no caches.
Publicly available, would meaningfully sharpen fund-level analysis without a budget.
Any one of these unlocks a large class of otherwise-blind reasoning. Plug in via a connector or env secret and the council picks it up automatically.
Terminal-grade coverage. Multi-thousand $/mo. Slot in when the research surface justifies the spend.
The Data requests panel on every run page is the council's live shopping list — exactly which series and documents the agents wished they had, ranked by how often they were requested across runs. That is your prioritised procurement roadmap, not a static wishlist.