我正在尝试使用Docker Composetesting一些需要HDFS的服务。 由于被testing的服务,名称节点和数据节点将全部运行在同一台物理机器上(开发笔记本电脑),所以通过只运行一个数据节点来降低内存使用率将是一件好事。 我正在使用这些泊坞窗图像 。
如果我运行一个名称节点和3个数据节点,所有按预期工作。 我试图通过在两个节点的hdfs-site.xml中设置这个节点来运行只有一个数据节点,并通过组合运行只有一个数据节点:
<property><name>dfs.replication</name><value>1</value></property>
这绝对是挑选这个,因为当它开始时,我在日志中看到这个:
blockmanagement.BlockManager: defaultReplication = 1 blockmanagement.BlockManager: maxReplication = 512 blockmanagement.BlockManager: minReplication = 1 blockmanagement.BlockManager: maxReplicationStreams = 2 blockmanagement.BlockManager: replicationRecheckInterval = 3000
第一次写入成功就好了。 对于第二次写,我得到了这个(在客户端应用程序;没有在hadoop方面logging):
java.io.IOException: Failed to replace a bad datanode on the existing pipeline due to no more good datanodes being available to try. (Nodes: current=[DatanodeInfoWithStorage[172.18.0.2:50010,DS-f97943bf-2cad-45e5-ae40-9ba947e54404,DISK]], original=[DatanodeInfoWithStorage[172.18.0.2:50010,DS-f97943bf-2cad-45e5-ae40-9ba947e54404,DISK]]). The current failed datanode replacement policy is DEFAULT, and a client may configure this via 'dfs.client.block.write.replace-datanode-on-failure.policy' in its configuration. at org.apache.hadoop.hdfs.DFSOutputStream$DataStreamer.findNewDatanode(DFSOutputStream.java:929) at org.apache.hadoop.hdfs.DFSOutputStream$DataStreamer.addDatanode2ExistingPipeline(DFSOutputStream.java:992) at org.apache.hadoop.hdfs.DFSOutputStream$DataStreamer.setupPipelineForAppendOrRecovery(DFSOutputStream.java:1160) at org.apache.hadoop.hdfs.DFSOutputStream$DataStreamer.run(DFSOutputStream.java:455)
之后的每次写入都会在客户端和HDFS上发生此错误:
Failed to APPEND_FILE (whatever) for (client X) on 172.18.0.6 because this file lease is currently owned by (client Y) on 172.18.0.6
如果运行3个数据节点,这个问题神奇地消失。 有没有人有任何经验在Docker中运行一个名称节点和一个数据节点? 我可怜的小笔记本电脑无法处理3个数据节点的功率水平。
编辑:我在这里试过这个解决scheme 。 没有骰子。 现在我得到:
17:59:56 WARN hdfs.DFSClient: DataStreamer Exception org.apache.hadoop.ipc.RemoteException(java.lang.UnsupportedOperationException): This feature is disabled. Please refer to dfs.client.block.write.replace-datanode-on-failure.enable configuration property. at org.apache.hadoop.hdfs.protocol.datatransfer.ReplaceDatanodeOnFailure.checkEnabled(ReplaceDatanodeOnFailure.java:116) at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getAdditionalDatanode(FSNamesystem.java:3317) at org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.getAdditionalDatanode(NameNodeRpcServer.java:758) [...]
HDFS方面的更丰富的日志( name是namenode, data是datanode;由于docker-compose,日志的交错并不完全按时间顺序排列):
name | 16/10/03 18:03:43 INFO hdfs.StateChange: BLOCK* registerDatanode: from DatanodeRegistration(172.18.0.11:50010, datanodeUuid=8ad27f17-7a87-45cb-b782-981c2e7b6dc2, infoPort=50075, infoSecurePort=0, ipcPort=50020, storageInfo=lv=-56;cid=CID-22dd8c41-af12-41ad-81ef-832ebb10ec39;nsid=1117453574;c=0) storage 8ad27f17-7a87-45cb-b782-981c2e7b6dc2 name | 16/10/03 18:03:43 INFO blockmanagement.DatanodeDescriptor: Number of failed storage changes from 0 to 0 data | 16/10/03 18:03:43 INFO datanode.VolumeScanner: VolumeScanner(/hadoop/dfs/data, DS-593eb971-f0cc-4381-a2c7-0befbc4aa9e6): finished scanning block pool BP-1023406345-172.18.0.9-1475517812059 data | 16/10/03 18:03:43 INFO datanode.DataNode: Block pool Block pool BP-1023406345-172.18.0.9-1475517812059 (Datanode Uuid null) service to hadoop-nn1/172.18.0.9:8020 successfully registered with NN data | 16/10/03 18:03:43 INFO datanode.DataNode: For namenode hadoop-nn1/172.18.0.9:8020 using DELETEREPORT_INTERVAL of 300000 msec BLOCKREPORT_INTERVAL of 21600000msec CACHEREPORT_INTERVAL of 10000msec Initial delay: 0msec; heartBeatInterval=3000 name | 16/10/03 18:03:43 INFO net.NetworkTopology: Adding a new node: /default-rack/172.18.0.11:50010 name | 16/10/03 18:03:44 INFO blockmanagement.DatanodeDescriptor: Number of failed storage changes from 0 to 0 data | 16/10/03 18:03:44 INFO datanode.VolumeScanner: VolumeScanner(/hadoop/dfs/data, DS-593eb971-f0cc-4381-a2c7-0befbc4aa9e6): no suitable block pools found to scan. Waiting 1814399359 ms. data | 16/10/03 18:03:44 INFO datanode.DataNode: Namenode Block pool BP-1023406345-172.18.0.9-1475517812059 (Datanode Uuid 8ad27f17-7a87-45cb-b782-981c2e7b6dc2) service to hadoop-nn1/172.18.0.9:8020 trying to claim ACTIVE state with txid=1 name | 16/10/03 18:03:44 INFO blockmanagement.DatanodeDescriptor: Adding new storage ID DS-593eb971-f0cc-4381-a2c7-0befbc4aa9e6 for DN 172.18.0.11:50010 data | 16/10/03 18:03:44 INFO datanode.DataNode: Acknowledging ACTIVE Namenode Block pool BP-1023406345-172.18.0.9-1475517812059 (Datanode Uuid 8ad27f17-7a87-45cb-b782-981c2e7b6dc2) service to hadoop-nn1/172.18.0.9:8020 data | 16/10/03 18:03:44 INFO datanode.DataNode: Successfully sent block report 0x8b0c17676f1, containing 1 storage report(s), of which we sent 1. The reports had 0 total blocks and used 1 RPC(s). This took 16 msec to generate and 190 msecs for RPC and NN processing. Got back one command: FinalizeCommand/5. name | 16/10/03 18:03:44 INFO BlockStateChange: BLOCK* processReport: from storage DS-593eb971-f0cc-4381-a2c7-0befbc4aa9e6 node DatanodeRegistration(172.18.0.11:50010, datanodeUuid=8ad27f17-7a87-45cb-b782-981c2e7b6dc2, infoPort=50075, infoSecurePort=0, ipcPort=50020, storageInfo=lv=-56;cid=CID-22dd8c41-af12-41ad-81ef-832ebb10ec39;nsid=1117453574;c=0), blocks: 0, hasStaleStorage: false, processing time: 2 msecs name | 16/10/03 18:04:33 INFO hdfs.StateChange: DIR* completeFile: /XXX/appender/1475517840000/.write/172.18.0.6 is closed by DFSClient_NONMAPREDUCE_1250587730_30 data | 16/10/03 18:03:44 INFO datanode.DataNode: Got finalize command for block pool BP-1023406345-172.18.0.9-1475517812059 data | 16/10/03 18:04:34 INFO datanode.DataNode: Receiving BP-1023406345-172.18.0.9-1475517812059:blk_1073741825_1001 src: /172.18.0.6:39732 dest: /172.18.0.11:50010 data | 16/10/03 18:04:34 INFO DataNode.clienttrace: src: /172.18.0.6:39732, dest: /172.18.0.11:50010, bytes: 7421, op: HDFS_WRITE, cliID: DFSClient_NONMAPREDUCE_1250587730_30, offset: 0, srvID: 8ad27f17-7a87-45cb-b782-981c2e7b6dc2, blockid: BP-1023406345-172.18.0.9-1475517812059:blk_1073741825_1001, duration: 107663969 name | 16/10/03 18:04:33 INFO hdfs.StateChange: BLOCK* allocate blk_1073741825_1001{UCState=UNDER_CONSTRUCTION, truncateBlock=null, primaryNodeIndex=-1, replicas=[ReplicaUC[[DISK]DS-593eb971-f0cc-4381-a2c7-0befbc4aa9e6:NORMAL:172.18.0.11:50010|RBW]]} for /XXX/appender/1475517840000/172.18.0.6 name | 16/10/03 18:04:34 INFO namenode.FSNamesystem: BLOCK* blk_1073741825_1001{UCState=COMMITTED, truncateBlock=null, primaryNodeIndex=-1, replicas=[ReplicaUC[[DISK]DS-593eb971-f0cc-4381-a2c7-0befbc4aa9e6:NORMAL:172.18.0.11:50010|RBW]]} is not COMPLETE (ucState = COMMITTED, replication# = 0 < minimum = 1) in file /XXX/appender/1475517840000/172.18.0.6 name | 16/10/03 18:04:34 INFO BlockStateChange: BLOCK* addStoredBlock: blockMap updated: 172.18.0.11:50010 is added to blk_1073741825_1001{UCState=COMMITTED, truncateBlock=null, primaryNodeIndex=-1, replicas=[ReplicaUC[[DISK]DS-593eb971-f0cc-4381-a2c7-0befbc4aa9e6:NORMAL:172.18.0.11:50010|RBW]]} size 7421 name | 16/10/03 18:04:34 INFO hdfs.StateChange: DIR* completeFile: /XXX/appender/1475517840000/172.18.0.6 is closed by DFSClient_NONMAPREDUCE_1250587730_30 name | 16/10/03 18:04:45 INFO namenode.FSEditLog: Number of transactions: 14 Total time for transactions(ms): 21 Number of transactions batched in Syncs: 1 Number of syncs: 8 SyncTimes(ms): 17 name | 16/10/03 18:04:45 INFO hdfs.StateChange: DIR* completeFile: /XXX/appender/1475517840000/.write/172.18.0.6 is closed by DFSClient_NONMAPREDUCE_-1821674544_30 name | 16/10/03 18:04:48 WARN hdfs.StateChange: DIR* NameSystem.append: Failed to APPEND_FILE /XXX/appender/1475517840000/172.18.0.6 for DFSClient_NONMAPREDUCE_1129971636_30 on 172.18.0.6 because this file lease is currently owned by DFSClient_NONMAPREDUCE_-1821674544_30 on 172.18.0.6
hdfs.site.xml (名称节点):
<configuration> <property><name>dfs.namenode.name.dir</name><value>file:///hadoop/dfs/name</value></property> <property><name>dfs.replication</name><value>1</value></property> <property><name>dfs.namenode.rpc-bind-host</name><value>0.0.0.0</value></property> <property><name>dfs.namenode.servicerpc-bind-host</name><value>0.0.0.0</value></property> <property><name>dfs.namenode.http-bind-host</name><value>0.0.0.0</value></property> <property><name>dfs.namenode.https-bind-host</name><value>0.0.0.0</value></property> <property><name>dfs.client.use.datanode.hostname</name><value>true</value></property> <property><name>dfs.datanode.use.datanode.hostname</name><value>true</value></property> </configuration>
hdfs-site.xml (数据节点):
<configuration> <property><name>dfs.datanode.data.dir</name><value>file:///hadoop/dfs/data</value></property> <property><name>dfs.replication</name><value>1</value></property> <property><name>dfs.namenode.rpc-bind-host</name><value>0.0.0.0</value></property> <property><name>dfs.namenode.servicerpc-bind-host</name><value>0.0.0.0</value></property> <property><name>dfs.namenode.http-bind-host</name><value>0.0.0.0</value></property> <property><name>dfs.namenode.https-bind-host</name><value>0.0.0.0</value></property> <property><name>dfs.client.use.datanode.hostname</name><value>true</value></property> <property><name>dfs.datanode.use.datanode.hostname</name><value>true</value></property> </configuration>
我通过在客户端以及服务器中设置dfs.replication属性来修复它。 任何遇到这个问题的人都应该尝试一下。
对于好奇,这里是我最终使用的docker-compose文件(如果你需要在Docker中设置一个快速的hdfs,应该节省一些时间):
version: "2" networks: platform: {} services: hdatanode: image: "uhopper/hadoop-datanode" networks: platform: aliases: - "hdatanode" environment: CORE_CONF_fs_defaultFS: "hdfs://hadoop:8020" CLUSTER_NAME: "cluster1" HDFS_CONF_dfs_replication: "1" depends_on: - "hadoop" hadoop: image: "uhopper/hadoop-namenode" networks: platform: aliases: - "hadoop" ports: - "50070:50070" - "8020:8020" environment: CLUSTER_NAME: "cluster1" HDFS_CONF_dfs_replication: "1"
然后,在客户端设置这三个configuration属性,如果在dockernetworking内运行:
<property><name>fs.defaultFS</name><value>hdfs://hadoop:8020</value></property> <property><name>fs.hdfs.impl</name><value>org.apache.hadoop.hdfs.DistributedFileSystem</value></property> <property><name>dfs.replication</name><value>1</value></property>
如果在dockernetworking外部运行,但是您的docker在本地主机上,则需要将其更改为:
<property><name>fs.defaultFS</name><value>hdfs://localhost:8020</value></property> <property><name>fs.hdfs.impl</name><value>org.apache.hadoop.hdfs.DistributedFileSystem</value></property> <property><name>dfs.replication</name><value>1</value></property>
即时HDFS用于testing/ dev目的!