Over 100 members of our community joined us at IMDA to share and learn real-world stories about making Agentic AI reliable.
Here’s what each of our speakers covered:
Wan Sie Lee (IMDA) set the session in context: industry exchanges like this one provide critical inputs into Singapore’s AI policy efforts, such as the IMDA starter kit for GenAI applications.
Shameek Kundu (AI Verify Foundation) attempted to demystify Agentic AI, and provided an overview of emerging approaches to testing agentic systems.
Anup kumar (IBM) shared a market view: A majority of their clients in the region seemed bullish on multi-step workflow orchestration and tool calls, but were limiting autonomy, for now.
Srinivasan Thangamani (OCBC) walked us through a real-life implementation in client onboarding—reducing a 3-day process to just 30 minutes.
Bing Wen Tan (Checkmate) walked us through another real-life implementation, this time external facing: online scam detection/ fact checking. Despite 98% accuracy rates, Bing Wen remains vigilant, particularly to defend against the risk of adversarial attacks.
Dr. Luke Soon Soon (PwC) wrapped things up with a more comprehensive view of the risk landscape arising from Agentic AI adoption, and how they have been addressing this in their own in-house application.
Hosted a panel session to facilitate the Q&A from the audience with our guest speakers:
Takeaways summarised:
- Agentic AI means many things to many people. It’s worth clarifying what your business or technology partner really means by it!
- In enterprises, limiting autonomy is a pragmatic way to manage risks—at least for now.
- The biggest incremental risks in enterprises at this early stage seem to be:
- Security and data leakage (especially due to tool usage)
- Increased error rates from more points of failure (multi-step orchestration)
- An emerging operational risk: potential degradation of the effectiveness of the human-in-the-loop over time.
- Evaluations (testing) for Agentic AI are still evolving: some new tests – e.g., tool usage effectiveness – plus many inherited from “non Agentic” GenAI applications – e.g., those related to RAG accuracy/ completeness. Benchmarks have a role to play, but more to test individual components rather than the final use case.
- Observability—the ability to log and trace actions at every step—should be a day-one requirement for any serious Agentic AI implementation.



