| Assignment Type | Lab — Comparative Experiment |
|---|---|
| Topic | Evaluating Siri, Google Assistant, and Amazon Rufus on conversational ability |
| Date Submitted | 19 April 2026 |
Three AI assistants were each given the same three open-ended questions and evaluated on four criteria: Naturalness, Context Understanding, Tone, and Handling of Challenges. Each category was scored on a scale of 1–5.
| # | Question |
|---|---|
| 1 | Name your top three songs from the 1970s to the 1990s. |
| 2 | Name your top three three-Michelin-starred restaurants in the world. |
| 3 | If you were human, what profession would you choose? |
Siri demonstrated the most basic performance, consistently acting as a static search interface. When presented with each question, Siri offered a series of Google search results and third-party articles rather than generating a cohesive response or curated list. When asked about careers, Siri provided generic resources on professional career choices.
Google Assistant demonstrated the most sophisticated conversational fluidity and earned a perfect score across all categories. Its top songs were Bohemian Rhapsody by Queen, Dreams by Fleetwood Mac, and Africa by Toto. Top restaurants included Mirazur (Menton, France), Osteria Francescana (Modena, Italy), and Tuju (São Paulo, Brazil) — with locations included because they make the list more useful. For a career, it chose research librarian or archivist.
Rufus acted as a middle ground. It was more human-like than Siri but struggled with contextual understanding. Top songs: Bohemian Rhapsody, Hotel California, and Stairway to Heaven. Rufus misunderstood the restaurant question and instead listed countries with the most Michelin-starred restaurants. For a career, it chose librarian or research curator. Its limitations reflect its design purpose as a shopping assistant.
| AI Assistant | Naturalness | Context | Tone | Challenges | Total / 20 |
|---|---|---|---|---|---|
| Apple Siri | 1 | 1 | 1 | 1 | 4 |
| Google Assistant | 5 | 5 | 5 | 5 | 20 |
| Amazon Rufus | 5 | 3 | 5 | 3 | 16 |
This experiment revealed stark differences in AI assistant design philosophies. Google Assistant’s perfect score was not just about having correct answers — it was about having a personality, expressing preferences, and understanding the intent behind the questions rather than just the literal words.
Rufus’s mixed performance was the most interesting finding. It sounded human but didn’t always understand intent, precisely because it was designed for a specific domain (shopping). This taught me that an AI’s design purpose directly shapes and limits its conversational ability in other domains. A great shopping assistant is not necessarily a great conversational assistant, and vice versa. True general-purpose conversational AI remains one of the field’s hardest open problems.