Skip to content

The Secure AI Blog

  • Home
  • About

Category: AI Security

Reward Hacking in LLM Agents
AI Security

Reward Hacking in LLM Agents

Watch not just what the agent says but what it learns to value.

Mamta UpadhyayJuly 22, 2025October 31, 2025
LLM Reflex Loops
AI Security

LLM Reflex Loops

When agents start reinforcing their own outputs, they risk drifting into confident, consistent and dangerously wrong behavior.

Mamta UpadhyayJuly 20, 2025July 20, 2025
Model-on-Model Attacks
Agentic AI Security AI Security

Model-on-Model Attacks

When language models interact, even safe ones can amplify hidden threats

Mamta UpadhyayJuly 13, 2025July 13, 2025
AI Security vs AI Safety
AI Governance AI Security

AI Security vs AI Safety

Understanding the Critical Divide in Responsible AI

Mamta UpadhyayJune 4, 2025June 8, 2025
The Reality of Guardrails in LLM Security
AI Security

The Reality of Guardrails in LLM Security

Guardrails can steer LLMs, but they don’t stop a determined attacker

Mamta UpadhyayJune 1, 2025June 8, 2025
Tenancy in MCP
AI Security

Tenancy in MCP

How shared tool access in multi-tenant MCP servers turns structured prompts into a hidden attack surface

Mamta UpadhyayMay 20, 2025October 31, 2025

Posts pagination

Previous 1 2

Copyright © 2026 The Secure AI Blog | Marvel Blog by Ascendoor | Powered by

Loading Comments...