FreeBSD / FreeNAS / TrueNAS

OK - last Sunday for the third Sunday in a row in August : my TrueNAS server was offline / powered off / shutdown…

It’s a HP N40L - 16 GB ECC RAM and 4 x 4 TB HDD zpool in RAIDZ+1 (i.e. like RAID 5) which gives me about 11 TB of capacity… The system comes with Dual Core AMD Turion CPU - it boots off a single SATA SSD (256 GB)…

No SMART issues or alerts… I’ve no idea what’s causing it… There was a scheduled job to “scrub” the zpool weekly - but didn’t co-incide with each outage… Nothing in any logs to indicate why… I can’t even diagnose the exact time it went offline / powered down…

So - I’ve disabled the scrub task - and - I now have a scheduled task (it uses cron - but you can’t see the crontab in the FreeBSD shell) on Thursday morning at 2:00 am to reboot… Then I also have cronjobs on my Pi4 system that uses the TrueNAS storage to restart the daemon that uses the storage, and same thing with JellyFin - about 3 am…

That worked fine overnight last night (this morning).

So - I’ll have to check if it’s still up (TrueNAS) on Sunday morning… Fingers crossed…

I might then change the scheduled scrub task to enabled but run on maybe Friday morning before sparrowfart (Aussie-ism for “dawn”)… See how I go…

Damn - I really should invest in a spare power supply for this box… That’s the mostly like thing to fail next - touch wood : those 4 x 4 TB (WD Green - i.e. 5400 RPM) have been going since late 2019 with no issues (previously had 2 x Seagate 3TB drives fail in it within two years). I can lose one HDD and retain functionality… Annoyingly - that AI bubble also seems to have affected spinning platter HDD prices too! It shouldn’t have… I reckon it’s price gouging, raising prices across the board…

That is the clue… what happens on Sunday ( or Sat night) ?
I reckon power supply switching. You WA mob have
batteries … I bet they are juggling things at the weekend.

Yeah Nah - it’s only my NAS server - my Ubuntu PC stays up and online, my Pi4 and Pi3, my Dell Optiplex running Debian 13… The Ubuntu PC, Debian 13 PC, Pi3 are all powered from the same outlet as the NAS…

Sunday sympton : I go to access something on the NAS from another PC and it’s not mounted… The NAS is powered off!

Let’s see what happens next sunday!

Do you have any cron jobs scheduled for Sunday?

No… There was a ZFS scrub task - but I’ve disabled it - just in case that’s what’s causing the halt / freeze / panic or whatever…

If it’s good tomorrow - I might re-enable the scrub task - but make it AFTER the reboot cronjob on Thursday morning…

Update - my NAS was still up this morning…

Starting to look like that ZFS Scrub task…

I’ve looked through all the logs from last time it happened (30/08)…

It logged something just after midnight - then basically nothing till I powered it up again about 9:15 am

But it also logged a massive list of arp thingies - i.e. even while it was supposedly unavailable… in /var/log/system :

Aug 30 00:20:13 baphomet kernel: arp: 10.1.1.227 moved from 34:e6:d7:60:71:e9 to 34:02:86:7f:2f:d8 on bge0
Aug 30 00:22:29 baphomet kernel: arp: 10.1.1.227 moved from 34:02:86:7f:2f:d8 to 34:e6:d7:60:71:e9 on bge0
Aug 30 00:42:32 baphomet kernel: arp: 10.1.1.227 moved from 34:e6:d7:60:71:e9 to 34:02:86:7f:2f:d8 on bge0
Aug 30 01:24:34 baphomet kernel: arp: 10.1.1.227 moved from 34:e6:d7:60:71:e9 to 34:02:86:7f:2f:d8 on epair0b
Aug 30 02:48:39 baphomet kernel: arp: 10.1.1.227 moved from 34:e6:d7:60:71:e9 to 34:02:86:7f:2f:d8 on bge0
Aug 30 03:50:36 baphomet kernel: arp: 10.1.1.227 moved from 34:e6:d7:60:71:e9 to 34:02:86:7f:2f:d8 on bge0
Aug 30 03:52:34 baphomet kernel: arp: 10.1.1.227 moved from 34:02:86:7f:2f:d8 to 34:e6:d7:60:71:e9 on bge0
Aug 30 04:32:39 baphomet kernel: arp: 10.1.1.227 moved from 34:e6:d7:60:71:e9 to 34:02:86:7f:2f:d8 on bge0
Aug 30 04:34:35 baphomet kernel: arp: 10.1.1.227 moved from 34:02:86:7f:2f:d8 to 34:e6:d7:60:71:e9 on bge0
Aug 30 04:54:37 baphomet kernel: arp: 10.1.1.227 moved from 34:e6:d7:60:71:e9 to 34:02:86:7f:2f:d8 on epair0b
Aug 30 05:14:38 baphomet kernel: arp: 10.1.1.227 moved from 34:e6:d7:60:71:e9 to 34:02:86:7f:2f:d8 on epair0b
Aug 30 05:14:40 baphomet kernel: arp: 10.1.1.227 moved from 34:e6:d7:60:71:e9 to 34:02:86:7f:2f:d8 on bge0
Aug 30 05:16:36 baphomet kernel: arp: 10.1.1.227 moved from 34:02:86:7f:2f:d8 to 34:e6:d7:60:71:e9 on bge0
Aug 30 05:36:39 baphomet kernel: arp: 10.1.1.227 moved from 34:e6:d7:60:71:e9 to 34:02:86:7f:2f:d8 on epair0b
Aug 30 06:18:39 baphomet kernel: arp: 10.1.1.227 moved from 34:e6:d7:60:71:e9 to 34:02:86:7f:2f:d8 on bge0
Aug 30 08:09:09 baphomet kernel: arp: 10.1.1.127 moved from 34:02:86:7f:2f:d8 to 34:e6:d7:60:71:e9 on bge0
Aug 30 08:09:10 baphomet kernel: arp: 10.1.1.227 moved from 34:02:86:7f:2f:d8 to 34:e6:d7:60:71:e9 on bge0
Aug 30 08:54:50 baphomet kernel: arp: 10.1.1.127 moved from 34:02:86:7f:2f:d8 to 34:e6:d7:60:71:e9 on bge0
Aug 30 08:54:52 baphomet kernel: arp: 10.1.1.227 moved from 34:e6:d7:60:71:e9 to 34:02:86:7f:2f:d8 on epair0b
Aug 30 09:58:57 baphomet kernel: arp: 10.1.1.127 moved from 34:02:86:7f:2f:d8 to 34:e6:d7:60:71:e9 on bge0

So - part of FreeBSD was still running at least… I see these sort of things in the logs constantly - I don’t know what they mean, nor do I particularly care… I think it’s a spurious false flag, but interesting in that it still logs this shyte when the system appears dead (note : the NIC LED flashes on the front even when the whole thing’s powered off - only way to stop that LED flashing is rip the power cable out of it).

So - I’ve moved that ZFS Scrub to 03:05 am on Thursday - i.e. 1 hour after the scheduled reboot… Let’s see what happens on Thursday morning - I’ll try and force the VGA console display to respond if it’s not available on the network (SSH, NFS and ping) - I have tiny 7" 4:3 VGA monitor plugged into it…

The annoying thing about scheduled scrub tasks - it’s just that - you don’t get to see what command it’s going to run… Anyway - I might try and trigger the symptom manually - apparently I can just run :
zpool scrub $ZPOOL (as root obviously)

There are 2 NIC’s and it is constantly moving(?) between them.
Why is zfs dealing with network?

No - there’s only a single NIC - neither of those IP addresses is associated with anything or used by anything…

bge0: flags=8843<UP,BROADCAST,RUNNING,SIMPLEX,MULTICAST> metric 0 mtu 1500
	description: biggsy
	options=c019b<RXCSUM,TXCSUM,VLAN_MTU,VLAN_HWTAGGING,VLAN_HWCSUM,TSO4,VLAN_HWTSO,LINKSTATE>
	ether 44:1e:a1:3d:fb:0c
	inet 10.2.1.10 netmask 0xfffc0000 broadcast 10.3.255.255
	inet 10.1.1.10 netmask 0xfffc0000 broadcast 10.3.255.255
	media: Ethernet autoselect (1000baseT <full-duplex>)
	status: active
	nd6 options=9<PERFORMNUD,IFDISABLED>

Note : this is a side issue - I don’t >really care< about it - it logs those messages all the time… It’s not associated with the hang/crash/panic/freeze… Other than it still logs them during that hung state…

If there’s something I learned (the hard way) in my home office, it’s having an extra power supply. Actually, two. It’s rare when it matters, but when it does, it can be quite unpleasant.

And yes, I’ve noticed HDD prices going up. Everything, really. I have never hated big tech as much as I do right now. I think we can all detect where this is going. In a decade, there will be no personal computing, and we’ll be doing everything off the cloud, for a subscription fee. The only holdouts left will be us weirdos who are keeping our old hardware and yelling ‘from my cold dead hands’! To which I say, so be it. On the positive, all of this has gotten me back into hardware tinkering. I’m even thinking about delving in ham radio and getting my license.

Common tasks like wordprocessing maybe.

I cant see scientific computing going that way … too inflexible.