Attention, KV Cache & Efficient Inference
← Question BankSign in to bookmark Sign in to bookmark Sign in to bookmark Sign in to bookmark Sign in to bookmark Sign in to bookmark Sign in to bookmark Sign in to bookmark Sign in to bookmark Sign in to bookmark Sign in to bookmark Sign in to bookmark Sign in to bookmark Sign in to bookmark
Deep Learning & Generative AI
Generative AI & Large Language Models (LLMs) Interview Questions
Interview questions on Generative AI & Large Language Models (LLMs) Interview Questions.
15 questions
All subtopicsLLM Architectures & FoundationsTokenization & EmbeddingsAttention, KV Cache & Efficient InferenceLLM Training & Pre-trainingDecoding Strategies & Prompt EngineeringFine-Tuning & Model AdaptationRAG & Vector DatabasesAlignment & Preference OptimizationLLM Evaluation & BenchmarkingAgents, Tool Use & MemoryMultimodal Generative AILLM Safety & SecurityLLMOps & Production
Sign in to bookmark
Attention, KV Cache & Efficient Inference
Q2. What are the main steps involved in the attention mechanism?
Attention, KV Cache & Efficient Inference
Q3. What is the KV cache in transformer models?
Attention, KV Cache & Efficient Inference
Q4. What are the advantages and disadvantages of using a Key-Value (KV) cache in Transformer models?
Attention, KV Cache & Efficient Inference
Q5. What is model distillation, and how does it benefit LLMs?
Attention, KV Cache & Efficient Inference
Q6. How do transformers improve on traditional Seq2Seq models?
Attention, KV Cache & Efficient Inference
Q7. What is multi-head attention, and how does it enhance LLMs?
Attention, KV Cache & Efficient Inference
Q8. What is local (sparse) attention and how does it work?
Attention, KV Cache & Efficient Inference
Q9. What is Flash Attention?
Attention, KV Cache & Efficient Inference
Q10. How is the softmax function applied in attention mechanisms?
Attention, KV Cache & Efficient Inference
Q11. How does the dot product contribute to self-attention?
Attention, KV Cache & Efficient Inference
Q12. How are attention scores calculated in transformers?
Attention, KV Cache & Efficient Inference
Q13. How does Mixture of Experts(MoE) enhance LLM scalability?
Attention, KV Cache & Efficient Inference
Q14. How do transformers address the vanishing gradient problem?
Attention, KV Cache & Efficient Inference