Prompt Injection Attacks
Prompt injection attacks are a type | of security vulnerability | where malicious inputs are crafted | to manipulate or exploit AI models,
Tấn công Prompt Injection là một loại | lỗ hổng bảo mật | trong đó các đầu vào độc hại được tạo ra | để thao túng hoặc khai thác các mô hình AI,
like language models, | to produce unintended or harmful outputs. | These attacks involve injecting | deceptive or adversarial content
như các mô hình ngôn ngữ, | nhằm tạo ra các đầu ra không mong muốn hoặc có hại. | Các cuộc tấn công này bao gồm việc tiêm | nội dung lừa đảo hoặc đối nghịch
into the prompt to bypass filters, | extract confidential information, | or make the model respond | in ways it shouldn’t.
vào prompt để vượt qua các bộ lọc, | trích xuất thông tin bảo mật, | hoặc khiến mô hình phản hồi | theo những cách mà nó không nên làm.
For instance, a prompt injection | could trick a model | into revealing sensitive data | or generating inappropriate responses
Ví dụ, một cuộc tấn công Prompt Injection | có thể đánh lừa một mô hình | tiết lộ dữ liệu nhạy cảm | hoặc tạo ra các phản hồi không phù hợp
by altering its expected behavior.
bằng cách thay đổi hành vi dự kiến của nó.
Resources
- Prompt Injection in LLMs (article)
- What is a Prompt Injection Attack? (article)
References
- https://roadmap.sh/ai-engineer (Node: Prompt Injection Attacks)
← AI Safety and Ethics · AI Engineer Roadmap · Bias and Fairness →