Hadoop

Hadoop is an open-source framework and ecosystem of software tools developed for processing and storing large volumes of data. Maintained by the Apache Software Foundation, it is designed to help organisations manage large and complex datasets efficiently.

Effective big data management plays an important role in both the operational performance and long-term development of a business. Hadoop-based technologies allow valuable data to be processed and used in different ways. They also enable data-intensive workloads that require significant computing power to be managed simultaneously across distributed environments.

The Hadoop ecosystem consists of several components used in data storage, processing and analysis. Its core components include:

Hadoop Distributed File System (HDFS): HDFS is a distributed file system designed to store large datasets. Data is divided into smaller blocks and distributed across multiple nodes, providing high availability, scalability and fault tolerance.

MapReduce: MapReduce is a distributed programming model used to process large datasets in parallel. It divides processing tasks across multiple nodes and consists of stages such as data splitting, mapping, shuffling and reducing.

YARN, or Yet Another Resource Negotiator: YARN is responsible for managing computing resources and scheduling workloads within a Hadoop cluster. It enables different applications to run on the same cluster while ensuring that available resources are allocated efficiently. This makes it easier to monitor and manage workloads across the Hadoop environment.

Hadoop Common: Hadoop Common provides the shared libraries, utilities and services required by the other Hadoop modules. It contains the core resources and tools needed to support distributed data processing.

The origins of the Hadoop ecosystem can be traced back to research papers published by Google on distributed file systems and large-scale data processing. Hadoop was later developed at Yahoo! as a practical framework for managing web-scale data.

Today, Hadoop is widely used in areas such as big data analytics, data warehousing, machine learning and artificial intelligence. Its open-source structure, scalability and support from a large global community have contributed significantly to its widespread adoption.

Discover it in the dictionary

Track the digital heartbeat with Kriko

Subscribe to receive curated insights, news, and ideas shaping the digital landscape.