Access
From iVEC User Documenation
Main Page > Access
Incident: 0005
Date/Time: 22/11/2012 9:20 AM
iVEC Operations Staff: Ashley Chew
Summary
Lustre file system issues.
Issue resolved.
Report:
22/11/2012 9:20 am (When it was noticed) Lustre Nodes foss1 to foss4 has kernel crashed. From the logs, problems starting around 8:30pm on 21st Nov. foss1 Nov 21 20:34:17 foss1 multipathd: sdc: tur checker reports path is down Nov 21 20:34:17 foss1 kernel: scsi 7:0:0:2: rejecting I/O to dead device Nov 21 20:34:17 foss1 multipathd: sdd: tur checker reports path is down Nov 21 20:34:17 foss1 kernel: scsi 7:0:0:3: rejecting I/O to dead device Nov 21 20:34:17 foss1 multipathd: sde: tur checker reports path is down Nov 21 20:34:17 foss1 kernel: scsi 7:0:0:4: rejecting I/O to dead device Nov 21 20:34:17 foss1 multipathd: sdf: tur checker reports path is down Nov 21 20:34:18 foss1 OpenSM[5276]: SM port is up foss2 Nov 21 20:34:46 foss2 kernel: scsi 7:0:0:7: rejecting I/O to dead device Nov 21 20:34:46 foss2 kernel: scsi 7:0:0:8: rejecting I/O to dead device Nov 21 20:34:46 foss2 kernel: scsi 7:0:0:9: rejecting I/O to dead device Nov 21 20:34:46 foss2 multipathd: sdj: tur checker reports path is down Nov 21 20:34:46 foss2 multipathd: sdk: tur checker reports path is down Nov 21 20:34:51 foss2 kernel: scsi 7:0:0:5: rejecting I/O to dead device Nov 21 20:34:51 foss2 kernel: scsi 7:0:0:6: rejecting I/O to dead device Nov 21 20:34:51 foss2 multipathd: sdg: tur checker reports path is down Nov 21 20:34:51 foss2 kernel: scsi 7:0:0:7: rejecting I/O to dead device foss3 Nov 21 20:28:05 foss3 kernel: scsi 7:0:0:11: rejecting I/O to dead device Nov 21 20:28:05 foss3 multipathd: sdl: tur checker reports path is down Nov 21 20:28:05 foss3 kernel: scsi 7:0:0:12: rejecting I/O to dead device Nov 21 20:28:05 foss3 multipathd: sdm: tur checker reports path is down Nov 21 20:28:05 foss3 kernel: scsi 7:0:0:13: rejecting I/O to dead device Nov 21 20:28:05 foss3 multipathd: sdn: tur checker reports path is down Nov 21 20:28:05 foss3 kernel: scsi 7:0:0:14: rejecting I/O to dead device Nov 21 20:28:05 foss3 multipathd: sdo: tur checker reports path is down Nov 21 20:28:05 foss3 multipathd: sdp: tur checker reports path is down foss4 Nov 21 20:28:52 foss4 kernel: LustreError: 8007:0:(filter_io_26.c:705:filter_commitrw_write()) error starting transaction: rc = -30 Nov 21 20:28:52 foss4 kernel: LustreError: 7998:0:(filter_io.c:507:filter_preprw_read()) Skipped 1 previous similar message Nov 21 20:28:52 foss4 kernel: LustreError: 7991:0:(filter_io_26.c:705:filter_commitrw_write()) error starting transaction: rc = -30 Nov 21 20:28:52 foss4 kernel: LustreError: 7997:0:(filter_io_26.c:705:filter_commitrw_write()) error starting transaction: rc = -30 Only changes that was done was srp_daemon which was done on Maintenance day (20th Nov 2012) ie srp_daemon -c -a -o -i mlx4_0 -p 1 -e to run_srp_daemon -R 20 -T 10 -nce -i mlx4_0 -p 1 It was changed to handle if a controller rebooted. Reverting changes to all OSS nodes (foss1-foss8), reboot all OSS nodes and restart lustre Back Online around 10:10 am Loss of storage was due to the DDN controller controller suddenly rebooting IS16000 DDN Storage Unit Last login: Thu Nov 22 10:57:14 2012 from 10.0.210.1 IS16000 RAID[0]$ show controller 0 all ************************* * Controller(s) * ************************* Index: 0 OID: 0x38000000 Firmware Version: Release: 1.4.2 Source Version: 8732 SGI Fully Checked In: Yes Private Build: No Build Type: Production Build Date and Time: 2011-11-18-21:54:UTC Builder Username: root Builder Hostname: lorax Build for CPU Type: AMD-64-bit Hardware Version: 0000 State: RUNNING Local AP OID: 0x00000000 Memory Size: 0x0 Max Q of S ID: 0x0 Up Time: 14 Hours 33 Minutes 15 Seconds Last Event Sequence #: 0x119ef Crash Dump Enabled: TRUE Log Disk Enabled: TRUE RP Count: 0x2 Restart Pending: FALSE Name: A Controller: LOCAL (SECONDARY) Controller ID: 0x0001ff0806350000 Enclosure OID: 0x50000000 (Index 0) Universal LAN Address: 0x00000001ff080635 MIR Reason: None NTP Sync: other controller Total Controllers: 1 Thu Nov 22 10:57:45 2012 IS16000 RAID[0]$ show controller 1 all ************************* * Controller(s) * ************************* Index: 1 OID: 0x38000001 Firmware Version: Release: 1.4.2 Source Version: 8732 SGI Fully Checked In: Yes Private Build: No Build Type: Production Build Date and Time: 2011-11-18-21:54:UTC Builder Username: root Builder Hostname: lorax Build for CPU Type: AMD-64-bit Hardware Version: 0000 State: RUNNING Local AP OID: 0x00000000 Memory Size: 0x0 Max Q of S ID: 0x0 Up Time: 18 Days 14 Hours 38 Minutes 30 Seconds Last Event Sequence #: 0x10eb Crash Dump Enabled: TRUE Log Disk Enabled: TRUE RP Count: 0x2 Restart Pending: FALSE Name: B Controller: REMOTE (PRIMARY) Controller ID: 0x0001ff0806380000 Enclosure OID: 0x5000000b (Index 11) Universal LAN Address: 0x00000001ff080635 MIR Reason: None NTP Sync: other controller Total Controllers: 1 One of the controllers rebooted (Not the same controller as last time)
Action Items:
Back to log



