Question 1
Discuss the key differences between structured, unstructured, and semi-structured data with examples.
Explain Hadoop's architecture in detail with a diagram. How does the NameNode and DataNode work together in Hadoop?
Discuss the key differences between structured, unstructured, and semi-structured data with examples.
Explain Hadoop's architecture in detail with a diagram. How does the NameNode and DataNode work together in Hadoop?
What is a MapReduce function? Explain the purpose of the Map and Reduce phases.
Discuss the concept of Data Locality in Hadoop. Why is it significant?
Write a detailed note on HDFS (Hadoop Distributed File System) and explain its significance in Big Data processing.
Discuss how MapReduce handles failures during the execution of jobs.
Define NoSQL databases and explain how they differ from relational databases.
Compare and contrast different types of NoSQL databases: Key- Value, Document, Column-family, and Graph databases.
Describe how NoSQL databases handle scalability and high availability.
What is Sharding in NoSQL databases? How does it help in handling Big Data?
Describe the architecture of MongoDB. Discuss its data model and how it handles queries and indexing.
Explain the role of Apache Spark in Big Data Analytics. How is it different from Hadoop MapReduce?
Discuss Spark’s RDD (Resilient Distributed Dataset) and its features.
Write and explain a simple Spark application for word count using PySpark
What is in-memory processing? Why is it beneficial in Big Data analytics?
Explain Spark’s transformation and action operations with examples.
How is Spark used to perform Machine Learning tasks? Discuss Spark MLlib with an example of a classification algorithm.
What is Mahout? Explain its use in scalable machine learning with Big Data.
Explain the concept of data streaming. How does Spark Streaming process real-time data?
Explain how data mining techniques are applied in Big Data Analytics. Mention some of the challenges faced when mining large datasets.
Describe the use of Big Data in retail industries. How can companies benefit from Big Data Analytics in decision-making?
Explain data warehousing in the context of Big Data. How does Hive enable querying large datasets?
Circulate this solved paper with KaTeX formulas and 1-click AI step solvers to your batchmates on WhatsApp or Telegram.
Official Gujarat Technological University (GTU) examination paper and step-by-step solutions for Big Data Analytics (BDA) (Summer 2025, B.E. · Computer Engineering, Sem 7). Features complete 70-mark regular & remedial examination pattern, official marking distribution across all 5 questions, and direct 1-click official PDF download.
Transcribed for student exam preparation from Gujarat Technological University official examination archives. Questions, syllabus guidelines, and curriculum marking schemes remain the intellectual property of Gujarat Technological University.