Access

From iVEC User Documenation
Main Page > Access
Jump to: navigation, search

Incident: 0005

Date/Time: 22/11/2012 9:20 AM

iVEC Operations Staff: Ashley Chew

Summary

Lustre file system issues.

Issue resolved.

Report:

22/11/2012 9:20 am (When it was noticed)
Lustre Nodes foss1 to foss4 has kernel crashed.
From the logs, problems starting around 8:30pm on 21st Nov.
foss1
Nov 21 20:34:17 foss1 multipathd: sdc: tur checker reports path is down
Nov 21 20:34:17 foss1 kernel: scsi 7:0:0:2: rejecting I/O to dead device
Nov 21 20:34:17 foss1 multipathd: sdd: tur checker reports path is down
Nov 21 20:34:17 foss1 kernel: scsi 7:0:0:3: rejecting I/O to dead device
Nov 21 20:34:17 foss1 multipathd: sde: tur checker reports path is down
Nov 21 20:34:17 foss1 kernel: scsi 7:0:0:4: rejecting I/O to dead device
Nov 21 20:34:17 foss1 multipathd: sdf: tur checker reports path is down
Nov 21 20:34:18 foss1 OpenSM[5276]: SM port is up
foss2
Nov 21 20:34:46 foss2 kernel: scsi 7:0:0:7: rejecting I/O to dead device
Nov 21 20:34:46 foss2 kernel: scsi 7:0:0:8: rejecting I/O to dead device
Nov 21 20:34:46 foss2 kernel: scsi 7:0:0:9: rejecting I/O to dead device
Nov 21 20:34:46 foss2 multipathd: sdj: tur checker reports path is down
Nov 21 20:34:46 foss2 multipathd: sdk: tur checker reports path is down
Nov 21 20:34:51 foss2 kernel: scsi 7:0:0:5: rejecting I/O to dead device
Nov 21 20:34:51 foss2 kernel: scsi 7:0:0:6: rejecting I/O to dead device
Nov 21 20:34:51 foss2 multipathd: sdg: tur checker reports path is down
Nov 21 20:34:51 foss2 kernel: scsi 7:0:0:7: rejecting I/O to dead device
foss3
Nov 21 20:28:05 foss3 kernel: scsi 7:0:0:11: rejecting I/O to dead device
Nov 21 20:28:05 foss3 multipathd: sdl: tur checker reports path is down
Nov 21 20:28:05 foss3 kernel: scsi 7:0:0:12: rejecting I/O to dead device
Nov 21 20:28:05 foss3 multipathd: sdm: tur checker reports path is down
Nov 21 20:28:05 foss3 kernel: scsi 7:0:0:13: rejecting I/O to dead device
Nov 21 20:28:05 foss3 multipathd: sdn: tur checker reports path is down
Nov 21 20:28:05 foss3 kernel: scsi 7:0:0:14: rejecting I/O to dead device
Nov 21 20:28:05 foss3 multipathd: sdo: tur checker reports path is down
Nov 21 20:28:05 foss3 multipathd: sdp: tur checker reports path is down
foss4
Nov 21 20:28:52 foss4 kernel: LustreError: 8007:0:(filter_io_26.c:705:filter_commitrw_write()) error starting transaction: rc = -30
Nov 21 20:28:52 foss4 kernel: LustreError: 7998:0:(filter_io.c:507:filter_preprw_read()) Skipped 1 previous similar message
Nov 21 20:28:52 foss4 kernel: LustreError: 7991:0:(filter_io_26.c:705:filter_commitrw_write()) error starting transaction: rc = -30
Nov 21 20:28:52 foss4 kernel: LustreError: 7997:0:(filter_io_26.c:705:filter_commitrw_write()) error starting transaction: rc = -30
Only changes that was done was srp_daemon which was done on Maintenance day (20th Nov 2012)
ie 
srp_daemon -c -a -o -i mlx4_0 -p 1 -e
to
run_srp_daemon -R 20 -T 10 -nce -i mlx4_0 -p 1
It was changed to handle if a controller rebooted.
Reverting changes to all OSS nodes (foss1-foss8), reboot all OSS nodes and restart lustre
Back Online around 10:10 am
Loss of storage was due to the DDN controller controller suddenly rebooting
IS16000 DDN Storage Unit
Last login: Thu Nov 22 10:57:14 2012 from 10.0.210.1
IS16000 RAID[0]$ show controller 0 all
*************************
*     Controller(s)     *
*************************
Index:                  0
OID:                    0x38000000
Firmware Version:
  Release:              1.4.2
  Source Version:       8732  SGI
  Fully Checked In:     Yes
  Private Build:        No
  Build Type:           Production
  Build Date and Time:  2011-11-18-21:54:UTC
  Builder Username:     root
  Builder Hostname:     lorax
  Build for CPU Type:   AMD-64-bit
Hardware Version:       0000
State:                  RUNNING
Local AP OID:           0x00000000
Memory Size:            0x0
Max Q of S ID:          0x0
Up Time:                14 Hours 33 Minutes 15 Seconds
Last Event Sequence #:  0x119ef
Crash Dump Enabled:     TRUE
Log Disk Enabled:       TRUE
RP Count:               0x2
Restart Pending:        FALSE
Name:                   A
Controller:             LOCAL        (SECONDARY)
Controller ID:          0x0001ff0806350000
Enclosure OID:          0x50000000 (Index 0)
Universal LAN Address:  0x00000001ff080635
MIR Reason:             None
NTP Sync:               other controller
Total Controllers: 1
Thu Nov 22 10:57:45 2012
IS16000 RAID[0]$ show controller 1 all
*************************
*     Controller(s)     *
*************************
Index:                  1
OID:                    0x38000001
Firmware Version:
  Release:              1.4.2
  Source Version:       8732  SGI
  Fully Checked In:     Yes
  Private Build:        No
  Build Type:           Production
  Build Date and Time:  2011-11-18-21:54:UTC
  Builder Username:     root
  Builder Hostname:     lorax
  Build for CPU Type:   AMD-64-bit
Hardware Version:       0000
State:                  RUNNING
Local AP OID:           0x00000000
Memory Size:            0x0
Max Q of S ID:          0x0
Up Time:                18 Days 14 Hours 38 Minutes 30 Seconds
Last Event Sequence #:  0x10eb
Crash Dump Enabled:     TRUE
Log Disk Enabled:       TRUE
RP Count:               0x2
Restart Pending:        FALSE
Name:                   B
Controller:             REMOTE       (PRIMARY)
Controller ID:          0x0001ff0806380000
Enclosure OID:          0x5000000b (Index 11)
Universal LAN Address:  0x00000001ff080635
MIR Reason:             None
NTP Sync:               other controller
Total Controllers: 1
One of the controllers rebooted (Not the same controller as last time)

Action Items:

Back to log

AccessSubscribeUser Portal
Main PageSupercomputersData StorageVisualisation