Answers

distcpUpdated August 2026

Why does distcp report success and still leave me with checksum mismatches on S3?

Because HDFS block checksums and S3 ETags are computed differently, distcp's default validation can pass while the objects are not byte-identical. The fix depends on whether you are using S3A's multipart upload and what you set for checksum handling on the job.

Cloudera exitUpdated July 2026

CDH support ends this year. Do we lift-and-shift to CDP Public Cloud or leave the platform entirely?

Both are defensible; the deciding factors are usually the volume of Hive/Impala logic you would have to rewrite, how much of the estate is already Spark, and whether the data itself can move ahead of the compute.

MetadataUpdated August 2026

The data copied fine. Why is migrating the Hive metastore the part that keeps failing?

Because the metastore encodes locations, partitions and statistics that all assume the old cluster. Version drift between on-prem Hive and the cloud service, and keeping both metastores in step during a hybrid period, are where most projects lose weeks.

HDFS to object storageUpdated June 2026

What actually breaks when an application written for HDFS starts reading from S3 or ADLS?

Rename is no longer atomic, listing is eventually consistent in some stores, small files become expensive, and directory semantics are emulated rather than real. Each of these has a known mitigation, but they need to be applied deliberately rather than discovered in production.

VerificationUpdated August 2026

How do we prove to an auditor that every file arrived, unchanged, when the source kept changing during the migration?

A one-time copy cannot do it, because the source data changed after the copy began. You need either a frozen source or a mechanism that tracks changes continuously and can produce a consistent, reconciled state at cut-over.

Data MigratorUpdated August 2026

Does Data Migrator need to be installed on the Hadoop cluster, and what does it do to NameNode load?

Data Migrator runs on an edge node with access to HDFS, not on the data nodes. It listens to the NameNode's edit log rather than scanning the filesystem, so the steady-state overhead is low; the initial scan is where you should plan capacity.

Replication and DRUpdated August 2026

Our DR plan for the data lake is a nightly distcp to a second site. What RPO does that really give us, and what does continuous replication change?

A nightly copy gives you an RPO of up to 24 hours plus however long the copy takes to complete, and the copy is inconsistent if the source changed during it. Continuous replication moves that to seconds, but the honest trade-offs are network cost, how you handle deletes, and what "consistent" means for a set of files that were never written as a transaction.

Replication and DRUpdated July 2026

Can we replicate HDFS to object storage across regions and still meet a regulator's requirement that the copy is provably identical at any point in time?

Yes, but only if the replication mechanism records what it has applied and can reconcile against the source, rather than re-scanning after the fact. What auditors usually want to see is a reconciliation report at a named timestamp, which means the tool has to know which source events the target reflects.

AI-ready dataUpdated August 2026

The AI team wants our Hive tables in Databricks and watsonx.data as Iceberg, but the source keeps changing. How do we keep both targets consistent without freezing production?

Treat it as two problems: the data files, which can be replicated continuously, and the table metadata, which must be translated (Hive Metastore to Iceberg catalog) and kept in step. The failure mode is targets that drift because metadata lands before the files it points to, so the ordering of the two streams is the whole design.

AI-ready dataUpdated June 2026

What does "data gravity" actually cost an AI project, and how do you decide whether to move the data or move the compute?

Gravity shows up as egress bills, latency, and the number of copies you end up maintaining. Move compute when the workload is one-off and the data is large; move data when several platforms need it or when the source system can't take the extra load. The rule of thumb we use is to count the consumers.

Replication and DRUpdated July 2026

What's the practical difference between DR for a Hadoop cluster and DR for a data lake on object storage?

A Hadoop DR plan protects a cluster: NameNode metadata, block placement and the services running on it. A data lake on S3 or ADLS has none of those, so the risk moves to the catalog, table metadata and the replication lag between regions or clouds. That changes what you test, and what "failover" even means, because there is no second cluster to fail over to.