Prompt Injection Attacks

Prompt Injection Attacks

Prompt injection attacks are a type | of security vulnerability | where malicious inputs are crafted | to manipulate or exploit AI models,
Tấn công Prompt Injection là một loại | lỗ hổng bảo mật | trong đó các đầu vào độc hại được tạo ra | để thao túng hoặc khai thác các mô hình AI,

like language models, | to produce unintended or harmful outputs. | These attacks involve injecting | deceptive or adversarial content
như các mô hình ngôn ngữ, | nhằm tạo ra các đầu ra không mong muốn hoặc có hại. | Các cuộc tấn công này bao gồm việc tiêm | nội dung lừa đảo hoặc đối nghịch

into the prompt to bypass filters, | extract confidential information, | or make the model respond | in ways it shouldn’t.
vào prompt để vượt qua các bộ lọc, | trích xuất thông tin bảo mật, | hoặc khiến mô hình phản hồi | theo những cách mà nó không nên làm.

For instance, a prompt injection | could trick a model | into revealing sensitive data | or generating inappropriate responses
Ví dụ, một cuộc tấn công Prompt Injection | có thể đánh lừa một mô hình | tiết lộ dữ liệu nhạy cảm | hoặc tạo ra các phản hồi không phù hợp

by altering its expected behavior.
bằng cách thay đổi hành vi dự kiến của nó.

Resources

References


← AI Safety and Ethics · AI Engineer Roadmap · Bias and Fairness →