Under the hood
All three are massive neural networks predicting next tokens. The architecture differences matter much less than people think. Same family, different siblings.
Where they actually differ
Training data composition. Reinforcement learning from human feedback shapes ‘personality.’ Tool-use plumbing differs (Gemini’s tied to Google search, Claude has Artifacts, GPT has DALL-E + Code Interpreter).
Practical
Try the same task in two of them. Pick whichever feels right. Switching costs are nearly zero. The real lock-in is your prompt library, not the model.