LLM Evals or Unit Tests? — Why Your CI Pipeline Is Lying
Discover why treating LLM evaluations like unit tests creates false security in CI/CD pipelines. Learn regression-testing discipline and metric validation...
The common belief that crafting perfect prompts is the most valuable AI skill has gained significant traction. Many believe that with the right prompts, language models (LLMs) can produce accurate and relevant responses. However, this myth overlooks the temporary nature of prompt engineering as a crucial skill.
While prompt engineering does have an impact on LLM performance, it is not a permanent skill. The reality is that LLMs are highly sensitive to prompt variations. For instance, negative sentiment in prompts reduces factual accuracy by approximately 8.4%, while positive sentiment reduces it by about 2.8%. Neutral prompts, on the other hand, yield the most factually accurate responses thebigdataguy.substack.com.
This sensitivity to prompts might lead one to believe that prompt engineering is a critical skill. However, as LLMs continue to evolve, the need for manual prompt optimization will diminish. Future advancements in AI will likely focus on developing more robust models that can understand context and intent, reducing the reliance on carefully crafted prompts. While prompt engineering may be a valuable skill in the short term, it is not a permanent one. As AI technology advances, the focus will shift from prompt engineering to more sophisticated and sustainable approaches to interacting with language models.
The increasing reliance on large language models (LLMs) has brought attention to the issue of hallucinations, where models generate factually incorrect or nonsensical responses. While prompt engineering has been touted as a solution to this problem, recent studies suggest that it is not a panacea. In fact, a 2025 study by Zep AI found that even with optimized prompts, hallucinations persist at an alarming rate.
One approach that has shown promise in reducing hallucinations is chain-of-thought prompting. This technique involves designing prompts that encourage the model to generate a step-by-step reasoning process, mimicking human thought. According to a study by Zep AI, chain-of-thought prompting reduced hallucinations from 38.3% to 18.1% Zep AI study. While this represents a significant improvement, it also highlights that prompt engineering alone is not enough to eliminate hallucinations.
The persistence of hallucinations, even with advanced prompting techniques, underscores the need for a deeper understanding of the underlying issues. It suggests that hallucinations are not solely a problem of input or prompt design but are also a symptom of more fundamental limitations in current LLM architectures.
The limitations of prompt engineering as a solution to hallucinations have significant implications for the development of reliable and trustworthy AI systems. As researchers and practitioners continue to explore new techniques for mitigating hallucinations, it is clear that a more comprehensive approach is needed—one that addresses both the input side (through context engineering and prompt optimization) and the model architecture itself. While chain-of-thought prompting represents a step in the right direction, it is crucial to recognize its limitations and the ongoing need for research into more robust solutions to the problem of hallucinations in LLMs. By understanding the strengths and weaknesses of current techniques, developers and researchers can work towards creating more reliable and accurate AI systems.
The field of AI interaction is undergoing a significant shift, one that moves beyond the traditional focus on perfecting prompts. This emerging paradigm, known as context engineering, is gaining traction among industry leaders and researchers. Context engineering, endorsed by Shopify CEO Tobi Lütke and AI researcher Andrej Karpathy, reframes development from crafting perfect instructions to dynamically curating what information enters the model's limited attention window at each step tao-hpu-medium.
The myth that prompt engineering alone can achieve optimal AI performance is being challenged by the emergence of context engineering. Prompt engineering, which focuses on crafting the perfect input to elicit a desired response from an AI model, has limitations. It assumes that the AI's performance is solely dependent on the quality of the prompt. However, this approach overlooks the importance of the information that the model has access to during inference.
Context engineering addresses this limitation by focusing on curating the input information that the model receives at each step. This approach recognizes that the model's performance is not just about the prompt, but also about the context in which the prompt is evaluated.
The rise of large language models (LLMs) has led to a surge in interest in prompt engineering, with many considering it a key skill for AI practitioners. However, as the industry moves towards more standardized interfaces, the need for hand-crafted prompts is diminishing. One such standardization effort is the Model Context Protocol (MCP), which aims to provide a unified interface for agents to query tools, databases, and external systems.
The Model Context Protocol (MCP) is designed to simplify interactions between agents and external systems, reducing the need for hand-crafted prompts. With over 10,000 public MCP servers deployed by late 2025 pub.towardsai.net, MCP is becoming a widely adopted standard for agent tool use.
MCP differs from traditional API integration in its focus on providing a standardized interface for agents to query tools and databases. This allows agents to interact with external systems in a more flexible and efficient manner, reducing the need for custom prompts.
By providing a standardized interface, MCP is making prompt engineering obsolete. As the industry moves towards more standardized interfaces, the need for hand-crafted prompts will continue to diminish, allowing agents to interact with external systems in a more efficient and effective manner.
The concept of GraphRAG has gained significant attention in recent times, particularly in the context of large language models (LLMs) and their applications. One key area where GraphRAG has shown promise is in improving query resolution efficiency and reducing the need for human intervention.
The idea that prompt crafting is the key to achieving high performance with LLMs has been a prevailing myth. However, evidence suggests that architectural improvements, such as those provided by GraphRAG, can have a much more substantial impact. According to a Fortune 500 enterprise case study, GraphRAG achieved 72% faster query resolution and a 40% reduction in SME (Subject Matter Expert) escalations [brlikhon.engineer blog].
This reality check is essential for developers and organizations looking to deploy LLMs effectively. Rather than focusing solely on prompt engineering, they should consider the underlying architecture and how it can be optimized for their specific use cases.
The rapidly evolving landscape of large language models (LLMs) has led to a surge in demand for skilled developers who can effectively integrate and leverage these technologies. However, with 95% of organizations running multiple LLM platforms simultaneously brlikhon.engineer, the focus has shifted from mere prompt engineering to more strategic skills. Indian developers, in particular, need to adapt their skill sets to stay competitive in this dynamic market.
The myth that prompt engineering is the primary skill for AI developers has been debunked by the reality of the market. With Anthropic, OpenAI, and Google Gemini vying for market share, their enterprise adoption rates have shown significant shifts. Anthropic captured 32% of enterprise market share in 2024, overtaking OpenAI's 25%, while Google Gemini grew from 7% to 21% enterprise adoption in just two years brlikhon.engineer. This volatility underscores the need for developers to focus on skills that are transferable across platforms.
Indian developers should prioritize skills in context engineering, Model Context Protocol (MCP), and GraphRAG. Context engineering, for instance, enables developers to design more effective prompts that can be used across different LLM platforms, thereby reducing the overhead of platform-specific customization.
Prompt engineering is a temporary interface, not a lasting skill. While it can reduce hallucinations—chain-of-thought prompting lowered them from 38.3% to 18.1% in a 2025 Zep AI study—LLMs remain highly sensitive to prompt variations. Negative sentiment in prompts reduces factual accuracy by ~8.4%, and positive sentiment by ~2.8%, making neutral prompts the most reliable. The industry is shifting toward context engineering and standardized protocols like MCP, which dynamically curate what enters the model's attention window, making manual prompt crafting less relevant for production systems.
Context engineering reframes AI interaction from crafting perfect instructions to dynamically curating the information fed into the model's limited attention window at each step. Endorsed by Shopify CEO Tobi Lütke and AI researcher Andrej Karpathy in mid-2025, this paradigm focuses on structuring and selecting relevant data—such as retrieved documents, tool outputs, or user history—rather than optimizing a static prompt string. Unlike prompt engineering, which relies on trial-and-error wording, context engineering leverages retrieval, routing, and memory systems to consistently improve factual accuracy and reduce hallucinations.
MCP provides a standardized interface for agents to query tools, databases, and external systems, eliminating the need for hand-crafted prompts to guide each interaction. By late 2025, over 10,000 public MCP servers were deployed, enabling agents to autonomously fetch context and execute actions through a uniform protocol. This shift means developers no longer need to engineer prompts for every API call or data source; instead, they focus on building robust context pipelines and agent orchestration logic.
Indian developers should prioritize context engineering, GraphRAG, and multi-platform integration. For example, GraphRAG implementations have achieved 72% faster query resolution and a 40% reduction in SME escalations in Fortune 500 enterprises. Additionally, with 95% of organizations running multiple LLM platforms simultaneously—and Anthropic capturing 32% enterprise market share in 2024, overtaking OpenAI's 25%—developers must learn to build portable, protocol-based systems (like MCP) rather than mastering prompt tricks for a single model.
Discover why treating LLM evaluations like unit tests creates false security in CI/CD pipelines. Learn regression-testing discipline and metric validation...
Understand how GraphQL resolvers fetch data, the N+1 problem's real impact, and how DataLoader with batching and caching solves it in production.
Learn to build production-grade autonomous AI agents from scratch. This 5-stage roadmap covers orchestration, reliability, and Python code for Indian...


