| IJ59717 |
Medium Importance |
Copy (cp command) of files from a .snapshot directory (subdirectories of export) using NFSv4.2 is not completed as 'cp' command hangs.
(show details)
| Symptom |
cp command will not be completed (block/hang is observed). |
| Environment |
Linux Only |
| Trigger |
Copy (cp command) files from a .snapshot directory (subdirectories ofexport) using NFSv4.2 mount |
| Workaround |
Mount the export containing .snapshot using NFSv4.0 or NFSv4.1 instead of NFSv4.2. This issue is not seen with NFSv4.0 and NFSv4.1 |
|
6.0.1.2 |
NFS |
| IJ59964 |
Suggested |
A script run by mmbuildgpl requires the which command. On RHEL10 and SLES16 that is provided by the which rpm package, that might not be installed on the system.
(show details)
| Symptom |
Error output/message |
| Environment |
ALL Linux OS environments |
| Trigger |
mmbuildgpl fails on RHEL10 and SLES16 when the which rpm package is not installed. |
| Workaround |
Manually install the which package. |
|
6.0.1.2 |
All Scale Users |
| IJ59713 |
Suggested |
GPFS does not allow the following together in one ACE: FileInherit flag AND DirInherit flag AND (Write perm XOR Append perm). This check was not previously enforced when an external NFSv4 ACL was converted to our internal representation. These entries can cause issues when working with the ACL on non-NFS file systems.
(show details)
| Symptom |
Unexpected Results/Behavior |
| Environment |
ALL Linux OS and AIX environments |
| Trigger |
Setting ACLs with nfs4_setfacl on NFS mount |
| Workaround |
Do not set ACLs with FileInherit flag AND DirInherit flag AND (Write perm XOR Append perm) |
|
6.0.1.2 |
All Scale Users |
| IJ59716 |
High Importance
|
Signal 11 error while performing rolling IO node reboot - openInactiveVdiskNsds()
(show details)
| Symptom |
Abend/Crash |
| Environment |
Linux Only |
| Trigger |
A remount Failure or unmounts on a busy system. |
| Workaround |
None |
|
6.0.1.2 |
GNR |
| IJ59714 |
High Importance
|
When running mmbackup incremental backup or update operations, mmbackup validates the number of files processed by comparing the candidate file count against the sum of files reported as backed up, updated, failed, or excluded by the IBM Storage Protect client. If the IBM Storage Protect client determines that a file's content and metadata are unchanged since the last backup, it does not count the file in any of those categories. Because mmbackup was not accounting for these unchanged files, it would incorrectly emit "[W] Some files missing during backup" or "[W] Some files missing during update" warnings even when all files were processed successfully.
IBM Storage Protect client version 8.2.2 introduced the -inclUnchangedCount option for incremental commands, which causes the client to report the count of unchanged files explicitly in its output. This fix enables mmbackup to detect when client version 8.2.2 or later is in use, automatically pass the -inclUnchangedCount option, and include the resulting unchanged count in its file count validation, eliminating the false warnings.
(show details)
| Symptom |
Error output/message |
| Environment |
ALL platforms that support mmbackup |
| Trigger |
This problem affects customers running mmbackup in incremental backup or update mode where one or more files are inspected by the IBM Storage Protect client but determined to be unchanged from the client's perspective (i.e., the file content and metadata have not changed since the last backup). All of the following conditions must be present for the spurious warning to occur:
- mmbackup is run in incremental backup or update mode
- One or more files inspected by the IBM Storage Protect client are considered unchanged by the client and are therefore not counted in the backed up, updated, failed, or excluded totals
- IBM Storage Protect client version prior to 8.2.2 is in use, or version 8.2.2 or later is in use without this fix applied to mmbackup |
| Workaround |
None |
|
6.0.1.2 |
mmbackup |
| IJ59613 |
HIPER |
pwritev2 is a system call available in all Linux distributions currently supported with IBM Storage Scale. That system call allows applications to specify the flags RWF_SYNC, RWF_DSYNC, RWF_APPEND and RWF_NOAPPEND. There are other flags that are out of scope of this discussion. The problem is that Storage Scale handled none of these flags and did not provide a notification back to the application that these are unsupported.
For the RWF_SYNC and RWF_DSYNC flag the result is that the data is not immediately synced to disk, even though this is required. The Linux kernel NFS server also uses these flags for stable NFS writes, those are affected in the same way in that the data is immediately synced to disk, even though stable NFS writes make that guarantee.
For the RWF_APPEND and RWF_NOAPPEND flags, IBM Storage Scale did not take action on those either, possible resulting in data being places in the wrong place in the file (in the middle of the file vs. at the end).
Depending on the Linux distro version, not all of these flags are available, e.g. RHEL8 does not include support for the RWF_NOAPPEND and RWF_APPEND flags and RHEL9 does not include support for the RWF_NOAPPEND flag. But for the flags that are supported on each Linux distro, the described problems exist.
(show details)
| Symptom |
Data loss |
| Environment |
ALL Linux OS environments |
| Trigger |
An application issues the pwritev2 system call with one of the RWF_SYNC, RWF_DSYNC, RWF_APPEND or RWF_NOAPPEND flag. For the flags supported by the current Linux distribution, the behavior of Storage Scale is not correct. Stable writes issues to the Linux NFS server on top of GPFS file system are also not correctly synced to disk. |
| Workaround |
The only workaround available is to not rely on these flags and issue explicit fsync calls and only rely on the open() O_APPEND flag. Since this might require changes to application code, these workarounds might not be available in most cases. |
|
6.0.1.2 |
All Scale Users |
| IJ59719 |
Suggested |
The mmfsck redirects the daemon command to the fs manager node via remote shell (e.g., ssh). It uses the hostname returned by the GPFS daemon. It first attempts to convert the daemon returned names to the reliable hostname stored in the SDR. If the name is unknown, mmfsck falls back to using the unknown name from the daemon for ssh. This becomes a problem when that name is not resolvable in DNS or not properly configured for promptless ssh/scp.
This problem only occurs only on clusters with different subdomains.
(show details)
| Symptom |
Command failure |
| Environment |
All |
| Trigger |
When run mmfsck command on cluster that have un-resolve partial hostnames (name with common domain removed but longer than short hostname). |
| Workaround |
Ensure that the partial names (with common domain removed) are resolvable |
|
6.0.1.2 |
Admin |
| IJ59720 |
Medium Importance |
Deleting dependent fileset in AFM for ro, lu, sw, iw modes can cause crash.
(show details)
| Symptom |
Crash observed on reuse of dependent fileset inode. |
| Environment |
All OS environments |
| Trigger |
Deletion of dependent fileset on AFM |
| Workaround |
Manually delete/unlink dependent fileset path on home first before deleting on cache cluster. |
|
6.0.1.2 |
AFM |
| IJ59768 |
Suggested |
Deleting SW dependent fileset on cache leaves files & directories under it at home.
(show details)
| Symptom |
Files availability on home cluster even after deletion of SW dependent fileset. |
| Environment |
All OS environments |
| Trigger |
Delete SW dependent fileset on cache cluster |
| Workaround |
Manual cleanup of left over dependent SW fileset paths on home cluster |
|
6.0.1.2 |
AFM |
| IJ59770 |
Medium Importance |
Hard link files are not cleaned or deleted when removed from home.
(show details)
| Symptom |
Hard link files keeps on growing and does not get deleted by async prefetch |
| Environment |
All OS environments |
| Trigger |
Async prefetch with hard link creation/deletion from source fileset |
| Workaround |
Manually delete the hardlink files in cache using mmafmlocal rm |
|
6.0.1.2 |
AFM |
IJ59690 |
Critical |
GPFS daemon could fail unexpectedly with assert: exp(fromNode != regP->owner). This could happen after a node failure or quorum loss.
(show details)
| Symptom |
Abend/Crash |
| Environment |
ALL Operating System environments |
| Trigger |
GPFS daemon failure, quorum loss or node expel |
| Workaround |
None |
|
6.0.1.2 |
All Scale Users |
| IJ59616 |
HIPER |
A dmapi enabled file system cannot be mounted on a remote cluster. The mount command will fail and the cluster manager node will log a mtsec error:
The mmfs.log on the client node reports a mount error:
2026-08-11_13:42:43.098-0700: [I] Command: mount fs1
2026-08-11_13:42:44.412-0700: [X] File System fs1 unmounted by the system with return code 13, reason code 0, at line 2643 in /project/sprelgpfs601/build/rgpfs6012623a/src/avs/fs/mmfs/ts/fs/mount.C
2026-08-11_13:42:44.412-0700: Permission denied
2026-08-11_13:42:44.412-0700: [W] Command: err 13: mount fs1
2026-08-11_13:42:44.412-0700: Permission denied
And the cluster manager node in the storage cluster logs a rejected RPC call:
2026-08-11_13:43:09.277-0700: [W] Access denied due to failed multitenancy authorization: RPC service 000E0001, msg 'ccMsgDmGetAllSessions', msg_id 222, from 10.21.102.186 home cluster name: storage-cluster-node-1.local err 13.
2026-08-11_13:43:09.603-0700: [E] File system fs1 unmounted by node 10.21.102.186 (remote-cluster-node-1.local in remote-cluster.local
(show details)
| Symptom |
Error output/message |
| Environment |
ALL Operating System environments |
| Trigger |
Simply mounting as dmapi enabled file system on a remote cluster with Storage Scale 6.0.1 or 6.0.1.1 will hit this problem. Describe the conditions necessary for the problem to occur in a language that the customer will understand. |
| Workaround |
The problem is due to a failed multitenancy check for the incoming RPC call to the storage cluster. Disabling this check will work around the problem:
echo 999 | mmchconfig mtsecCheckEnabled=0 -i
Note that this reduces the protection against attacks from untrusted remote clusters to the storage cluster back to the 6.0.0 level.
Once the fix is installed, revert the config change to re-enable the security feature:
echo 999 | /usr/lpp/mmfs/bin/mmchconfig mtsecCheckEnabled=DELETE -i |
|
6.0.1.2 |
DMAPI/HSM/TSM |
| IJ58285 |
High Importance
|
Kernel crash due to segFault if file is opened in cache and same file is changed at home.
(show details)
| Symptom |
Abend/Crash |
| Environment |
Linux Only |
| Trigger |
File is changed at home while it is opened by application in cache. |
| Workaround |
None |
|
6.0.1.2 |
AFM |
| IJ59929 |
High Importance
|
Read and write operations on FIFO (named pipe) files in a GPFS file system hang indefinitely on nodes running version 5.2.3.8 or later. Creating the FIFO succeeds and the file is visible, but any process that attempts to read from or write to it blocks forever.
(show details)
| Symptom |
Kernel panic/ crash |
| Environment |
All Linux OS environments |
| Trigger |
Any process that attempts to read from or write to fifo blocks forever |
| Workaround |
There is no workaround |
|
6.0.1.2 |
All Scale Users using FIFOs. |
| IJ58748 |
High Importance
|
When many threads perform concurrent block allocation, it is possible for same region to be selected for use leading to contention for lock on the region.
(show details)
| Symptom |
Performance Impact/Degradation |
| Environment |
ALL Operating System environments |
| Trigger |
Many threads perform block allocation concurrently |
| Workaround |
None |
|
6.0.1.2 |
All Scale Users |
| IJ58616 |
High Importance
|
When shutting down nodes, and incomplete rebuild, the FT is not correctly shown.
(show details)
| Symptom |
Unexpected Results/Behavior |
| Environment |
ALL Operating System environments |
| Trigger |
Decrease the number of spare disks in the DA, and turn down couple of nodes. |
| Workaround |
None |
|
6.0.1.2 |
ESS/GNR |
| IJ57903 |
High Importance
|
Unexpected log waiter could occur on directory operations such as file/directory create when parent directory has missing block.
(show details)
| Symptom |
Hang/Deadlock/Unresponsiveness/Long Waiters |
| Environment |
ALL Operating System environments |
| Trigger |
Corruption require directory block to be deleted by fsck |
| Workaround |
Move directory content to another directory |
|
6.0.1.2 |
All Scale Users |
| IJ59954 |
High Importance
|
The mmafmctl checkUncached command fails to report uncached files from dependent filesets.
(show details)
| Symptom |
Unexpected results |
| Environment |
Linux Only |
| Trigger |
Running mmafmctl checkUncached command on fileset or file system with dependent filesets |
| Workaround |
None |
|
6.0.1.2 |
AFM |
| IJ59951 |
High Importance
|
AFM replication messages may get stuck on the NFS mount when NFS server is not responding which may cause deadlock.
(show details)
| Symptom |
Deadlock |
| Environment |
Linux Only |
| Trigger |
Unresponsive NFS server when AFM requests are inflight. |
| Workaround |
None |
|
6.0.1.2 |
AFM |
| IJ59946 |
High Importance
|
Deadlock occurs while stopping the AFM fileset due to incorrect handler count verification.
(show details)
| Symptom |
Deadlock |
| Environment |
Linux Only |
| Trigger |
Stopping AFM fileset |
| Workaround |
None |
|
6.0.1.2 |
AFM |
IJ59956 |
Critical |
When opening file handles with the O_DIRECT flag and passing those into the copy_file_range system call, this will fail when running Storage Scale 6.0.1 PTF1.
(show details)
| Symptom |
IO error |
| Environment |
ALL Operating System environments |
| Trigger |
Open two file descriptors, at least one with the O_DIRECT flag. Use them for the copy_file_range system call. In Storage Scale 6.0.1 PTF1 the copy_file_range call will fail. |
| Workaround |
Do not open file handles with O_DIRECT when using the copy_file_range system call, or implement copies with read and write system calls. |
|
6.0.1.2 |
All Scale Users |
| IJ59950 |
High Importance
|
AFM migration fails to pull files that were recreated before the cutover when afmPtrashOpt=3 is set. This occurs due to incorrect conflict resolution in the AFM cache. This affects only 6.0.0.x releases.
(show details)
| Symptom |
Unexpected results |
| Environment |
Linux Only |
| Trigger |
AFM migration with afmPtrashOpt=3 value |
| Workaround |
Set config option afmPtrashOpt=0 |
|
6.0.1.2 |
AFM |
| IJ59948 |
High Importance
|
AFM gateway node crashes during the file system unmount with the assert "gpfs_s_put_super: s_inodes not empty" due to a bug in conflict resolution.
(show details)
| Symptom |
Crash |
| Environment |
Linux Only |
| Trigger |
AFM caching with conflict resolution |
| Workaround |
None |
|
6.0.1.2 |
AFM |
IJ59949 |
Critical |
In-memory data mismatch issue could occur with a NFS hard mount, where partially read data is first written locally and then cleared later due to an error caused by killing stuck mount requests. If an application reads the data (on a different thread, actual threads performing read may receive EIO error) between the time it is written locally and the time it is cleared, it may see incorrect data. However, the data on disk would still be correct.
(show details)
| Symptom |
Unexpected results |
| Environment |
Linux Only |
| Trigger |
AFM caching with unresponsive NFS server |
| Workaround |
None |
|
6.0.1.2 |
AFM |
| IJ59947 |
High Importance
|
AFM recovery incorrectly moves renamed directories to the .afmtrash directory at the home location. This happens when the contents of a directory are renamed to a different directory and the original directory itself is removed. During recovery, AFM incorrectly queues the Rmdir operation first, which may cause the entries to be moved to the .afmtrash directory at the home location.
(show details)
| Symptom |
Unexpected results |
| Environment |
Linux Only |
| Trigger |
AFM recovery with dependent Remove and Rename operations. |
| Workaround |
None |
|
6.0.1.2 |
AFM |
| IJ59952 |
Suggested |
mmafm log flooding can cause ENOSPACE issues on cmd operations
(show details)
| Symptom |
mmafm log file size grows much faster |
| Environment |
All OS environments |
| Trigger |
Large fileset sync using prefetch on newly created filesets. |
| Workaround |
Disable mmafm log, mmchconfig afmLogLevel=0 -i |
|
6.0.1.2 |
AFM |
| IJ59963 |
High Importance
|
While direct I/O workloads are running on a QoS-enabled filesystem, QoS statistics collection can trigger a NULL pointer dereference, resulting in a kernel panic.
(show details)
| Symptom |
Abend/Crash |
| Environment |
ALL Operating System environments. |
| Trigger |
Frequent QoS configuration updates and refresh operations during heavy direct I/O workloads. |
| Workaround |
None |
|
6.0.1.2 |
QoS |
| IJ59955 |
Suggested |
Hardlink will be broken if original file is deleted from cache.
(show details)
| Symptom |
operation stuck |
| Environment |
Linux Only |
| Trigger |
Operation gets stuck while reading the hardlinks. |
| Workaround |
None |
|
6.0.1.2 |
AFM |
| IJ59788 |
High Importance
|
queue drain slowed which caused mmshutdown to become slow.
(show details)
| Symptom |
Performance Impact |
| Environment |
Linux Only |
| Trigger |
create 100-200 filesets and trigger recovery on fileset |
| Workaround |
set afmMaxParallelRecoveries=5 |
|
6.0.1.2 |
AFM |
| IJ59789 |
Suggested |
This fix addresses an issue involving a race condition between two threads accessing the same file. While one thread was writing to the file and updating the size, another thread was reading the file size through a stat call concurrently, resulting in the stat returning a stale file size.
(show details)
| Symptom |
stat system call returns a stale file size |
| Environment |
ALL Operating System environments |
| Trigger |
This is a narrow race condition that is hit if a thread is doing fast cached writes to a file and another thread calls stat on that same file concurrently. |
| Workaround |
None |
|
6.0.1.2 |
All Scale Users |
| IJ59992 |
Critical |
GPFS internally queues mmap writeback requests and processes them for worker threads. Due to a timing issue, and incorrect handling of waiting threads, a situation can arise where no worker threads are picking up entries from this queue. This results in stuck mmap writeback.
(show details)
| Symptom |
Hang/Deadlock/Unresponsiveness/Long Waiters |
| Environment |
ALL Linux OS environments (except RHEL8 and except ppc64le) |
| Trigger |
This problem has been seen with heavy mmap workloads with Scale 6.0.1.0 and 6.0.1.1. RHEL8 is using the old semaphore based lockingl, so is not affected. ppc64le is also not affected. |
| Workaround |
While the problem is in the mmap code, the timing has been introduced with a change in low-level locking. This locking change can be disabled, to avoid the mmap problem:
Change the config for all nodes: mmchconfig kernelLockType=2
The one node at a time, restart GPFS to pick-up the config change:
mmshutdown
mmstartup |
|
6.0.1.2 |
All Scale Users |
| IJ60020 |
High Importance
|
IBM Storage Scale experiences data integrity issues when data in the Scale filesystem is accessed through a Linux loop device and the system is running a newer Linux kernel level. The issue was exposed by a Linux kernel change that modified how loop devices perform buffered I/O, requiring change in Storage Scale's handling of read and write requests for loop devices.
(show details)
| Symptom |
Silent data corruption, inconsistent checksums across mount/unmount cycles from loop-mounted filesystems. |
| Environment |
Linux Only (Linux kernel ≥ 6.15, SLES ≥ 6.4.0-150700.53.6-default, RHEL ≥ 5.14.0-597, Ubuntu ≥ 6.14.4) |
| Trigger |
Accessing iso files or data in other loop mounted files residing on an IBM Storage Scale filesystem via Linux loop devices on systems running Linux kernels with the updated loop buffered I/O mechanism. |
| Workaround |
None |
|
6.0.1.2 |
Linux Kernel Extension (VFS / file operations layer) |
| IJ60049 |
High Importance
|
mmafmcoskeys all delete can wipe out all COS buckets configured for keys.. sometimes 100s or even 1000s.
(show details)
| Symptom |
Unexpected Behavior |
| Environment |
All OS Environments |
| Trigger |
mmafmcoskeys all delete |
| Workaround |
Be very careful in not using the all keyword to delete mmafmcoskeys |
|
6.0.1.2 |
AFM |
| IJ60050 |
High Importance
|
If prefetch is run with a list file and --delete option on IW mode, it releases the list of inodes but also makes the parent directories of removed inodes as dirty preventing future prefetch (without delete) from fetching the removed inodes from the remote site.
(show details)
| Symptom |
Unexpected Behavior |
| Environment |
All OS Environments |
| Trigger |
mmafmctl prefetch with --list-file and --delete option on an IW/SW mode fileset. |
| Workaround |
None |
|
6.0.1.2 |
AFM |
| IJ60051 |
High Importance
|
RO mode configured with refresh intervals disabled can fail to prefetch anything although prefetch command is run every X mins like clockwork through the Async Prefetch tunable.
(show details)
| Symptom |
Unexpected Behavior |
| Environment |
Linux only (AFM Gateways) |
| Trigger |
Having a fileset with disabled refresh intervals and asyncPrefetchInterval enabled every X minutes. |
| Workaround |
Re-enable or shorten the refresh intervals. |
|
6.0.1.2 |
AFM |
| IJ58836 |
High Importance
|
Negative lookups are queued to gateway after cache drop on client
(show details)
| Symptom |
Unexpected Results |
| Environment |
Linux Only |
| Trigger |
run negative lookup (ls on non-existant file ) from non-gw node and clear cache from non-gw |
| Workaround |
None |
|
6.0.1.1 |
AFM |
| IJ58969 |
Suggested |
When a symlink is accessed on another node, then deleted on a different node and on the first node the Linux struct inode is reused, some memory might not be deallocated.
(show details)
| Symptom |
Unexpected Results/Behavior |
| Environment |
ALL Linux OS environments |
| Trigger |
The scenario is to have a symlink access it on node A. Then delete the symlink on a different node. When a different file is created and accessed on node A, some memory allocated when accessing the symlink |
| Workaround |
There is no workaround. |
|
6.0.1.1 |
All Scale Users |
| IJ58838 |
High Importance
|
mmsmb exportacl list command will list the output in SIDs instead of user/group names when there are more than 6000 unique SIDs in the system.
This is because the command uses rpcclient to resolve SIDs to user/group names, and rpcclient has a limit of 6000 SIDs that it can resolve at once.
When there are more than 6000 unique SIDs, rpcclient fails to resolve them, and mmsmb exportacl list will show the SIDs instead of user/group names.
Warning message is seen in the logs:
mmsmb exportacl list: Execution response: Usage: rpcclient [OPTION...] BINDING-STRING|HOST
(show details)
| Symptom |
Error output/message
Warning message is seen in the logs: mmsmb exportacl list: Execution response: Usage: rpcclient [OPTION...] BINDING-STRING|HOST |
| Environment |
Linux Only |
| Trigger |
Customer having more than 6000 unique SIDs in the system. |
| Workaround |
Customer can still see the SIDs in the output of mmsmb exportacl list command, but they will not be able to see the user/group names. |
|
6.0.1.1 |
CES |
| IJ58970 |
High Importance
|
In a heavy workload that opens and closes an excessive number of files and directories, a kernel leak may continuously grow over days and result in the depletion of the Windows non-paged pool, requiring a node reboot to restore it to a healthy state.
(show details)
| Symptom |
Unexpected Results/Behavior. |
| Environment |
Windows/x86_64 only |
| Trigger |
Heavy workload that opens and closes an excessive number of files and directories over an extended time. |
| Workaround |
None |
|
6.0.1.1 |
Windows |
| IJ58917 |
High Importance
|
AFM does not failover to another available NFS Server in the export map using afmFailoverMap casing the deadlock
(show details)
| Symptom |
Deadlock |
| Environment |
Linux Only |
| Trigger |
NFS failover with afmFailoverMap |
| Workaround |
None |
|
6.0.1.1 |
AFM |
| IJ58919 |
Medium Importance |
If a node is promoted to quorum node by "mmchnode --quorum" after one or more other cluster nodes are already expelledand "disablePersistExpelList=false" (the default), the expelled node list is not set up on the new quorum node, so if the new quorum node subsequently becomes the leader (cluster manager), the expelled nodes will be allowed to rejoin.
(show details)
| Symptom |
Nodes may unexpectedly rejoin the cluster. |
| Environment |
All |
| Trigger |
Newly-added quorum node and persistent expelled node(s) |
| Workaround |
None |
|
6.0.1.1 |
GPFS core |
| IJ58971 |
Suggested |
Various deficiencies in Multitenant Security RPC (MTSec) Filtering were not resolved in 6.0.1.0 GA (where the feature first appeared).
- mtSecPrintResetInterval and mtSecPrintLimitReset should limit the number of messages output to mmfs.log during an attack or a defect, but the filtering reset limits did not work properly.
- Changed to do rate limited on a per-node basis
- Fix incorrect aborting from bcast_takeover_query & doRfacCleanup
- Removed some duplicated messages
- Improved reporting for some NSD and related MTSec errors to show the offending sender node instead of NODE_NONE
(show details)
| Symptom |
Warning messages in mmfs.log |
| Environment |
All |
| Trigger |
Remote cluster storage clients generating impermissible RPCs. |
| Workaround |
N/A |
|
6.0.1.1 |
GPFS core |
| IJ58840 |
High Importance
|
Under certain workloads, the token manager nodes can experience a contention on the aTokeClassMutex, leading to performance degradation. All token revokes for the same token type on each token server has to acquire the same mutex. The symptom would mmdiag --waiters on the token manager nodes shows many short waiters for aTokenClassMutex.
(show details)
| Symptom |
Performance Impact/Degradation |
| Environment |
ALL Operating System environments |
| Trigger |
This is hit with a workload that triggers many token revokes. Having fewer token manager nodes might exacerbate the problem, since the mutex is local to each token server. |
| Workaround |
Modifying the workload to reduce token contention could avoid this problem. Also adding more manager nodes to spread the token manager load could reduce this problem. |
|
6.0.1.1 |
All Scale Users |
| IJ58920 |
High Importance
|
Revalidation Performance Degradation After a Fileset Is Converted to Primary Mode. This is due to Open requests being sent to the target/secondary.
(show details)
| Symptom |
Performance Impact |
| Environment |
Linux Only |
| Trigger |
This issue occurs when a fileset is converted to an AFM Primary mode fileset. |
| Workaround |
None |
|
6.0.1.1 |
AFM |
| IJ58972 |
Suggested |
When vlans are used on IP over IB devices with Vlans an event will be raised unless vlans are used on a bond.
(show details)
| Symptom |
Error output/message |
| Environment |
ALL Linux OS environments |
| Trigger |
When vlans are used on IP over IB devices with Vlans an event will be raised unless vlans are used on a bond. |
| Workaround |
None |
|
6.0.1.1 |
System Health |
| IJ58921 |
High Importance
|
Revalidation Performance Degradation After a Fileset Is Converted to Primary Mode. This is due to Open requests being sent to the target/secondary.
(show details)
| Symptom |
Performance Impact |
| Environment |
Linux Only |
| Trigger |
This issue occurs when a fileset is converted to an AFM Primary mode fileset. |
| Workaround |
None |
|
6.0.1.1 |
AFM |
| IJ58973 |
High Importance
|
When Call Home is configured to use a custom CA bundle (e.g. via the CURL_CA_BUNDLE environment variable set in the mmsysmon service configuration), upload operations that transmit data fail with the following SSL error:
curl: (60) SSL certificate problem: self-signed certificate in certificate chain
All other Call Home operations succeed, including connectivity tests performed by "mmcallhome test connection" succeed. Only the actual data upload path is affected.
(show details)
| Symptom |
- Error output/message
- Component Level Outage |
| Environment |
ALL Linux OS environments |
| Trigger |
All of the following conditions must be met:
1. IBM Storage Scale 5.2.2.1 or later is installed.
2. A custom CA bundle is configured for the mmsysmon service via the CURL_CA_BUNDLE environment variable (typically in /etc/systemd/system/mmsysmon.service.d/).
3. The environment uses a self-signed or enterprise-issued CA certificate that is not present in the system default CA store.
4. A Call Home upload operation is triggered (automatic daily/weekly gather-and-send, or a manual mmcallhome send). |
| Workaround |
On the node designated as the Call Home server, apply the following manual code change in
/usr/lpp/mmfs/lib/mmsysmon/callhome/Callhome.py:
@@ -173,6 +173,8 @@ class CallhomeJobPrototype(object):
"""Performs an upload attempt and returns its rc and stdout"""
self.logger.info("Running command: %s", " ".join(command)
CURL_ERROR_PREFIX = "curl: ("
+ fullenv = dict(os.environ)
+ fullenv.update({"https_proxy": CallhomeConfig().proxyCurlEnvParameter})
try:
proc = subprocess.Popen(
command,
@@ -181,7 +183,7 @@ class CallhomeJobPrototype(object):
stderr=subprocess.PIPE,
close_fds=True,
encoding="utf-8",
- env={"https_proxy": CallhomeConfig().proxyCurlEnvParameter},
+ env=fullenv,
)
lastTrailingLine = ""
# for preventing the same percentage being written in log file
Then restart the mmsysmon daemon:
mmsysmoncontrol restart |
|
6.0.1.1 |
Callhome |
| IJ58899 |
Medium Importance |
Inode used not decreasing with auto eviction & inode quota enabled
(show details)
| Symptom |
Slow or no progress in release of inodes with auto-eviction |
| Environment |
All OS environments |
| Trigger |
Auto eveiction is configured on fileset, emptyPtrashErrors.xxxx accumulate over time. |
| Workaround |
Manually analyze .ptrash & .afm directory and delete emptyPtrashErrors.xxxx files |
|
6.0.1.1 |
AFM |
| IJ58925 |
High Importance
|
empty directory and symlink mtime is not in sync with remote.
(show details)
| Symptom |
Unexpected Results |
| Environment |
Linux Only |
| Trigger |
run dmapi prefetch for empty dir or symlink. |
| Workaround |
run mmafmlocal rm -rf <emptyDir / symlink>
then run ls <emptyDir / symlink></emptyDir> |
|
6.0.1.1 |
AFM |
| IJ58843 |
Suggested |
mmbackup may incorrectly report successful completion when some files actually failed to backup. This occurs due to three specific issues: (1) The ANS4047E error from TSM (indicating a read error on a file that is skipped) is not included in the error filter list, causing mmbackup to miss this failure condition. (2) When dsmc reports "Total number of objects failed: <n>", mmbackup does not correctly count these failures. (3) mmbackup does not report "some files are missing" when only file updates (metadata changes) fail, even though backup operations for new or changed file content may have succeeded.
(show details)
| Symptom |
Error output/message |
| Environment |
All platforms that support mmbackup |
| Trigger |
This issue affects users running mmbackup operations when any of the following conditions occur:
1) Files encounter read errors during backup, generating ANS4047E errors from TSM. This can happen when:
- Files are being modified during the backup operation
- File permissions change during backup
- Files are deleted after being selected for backup but before being read
- I/O errors occur when reading file data
2) Protect client reports failed objects in its output with "Total number of objects failed: ", but these failures are not properly counted by mmbackup
3) During incremental backup operations where only file metadata needs to be updated (not file content), and these update operations fail |
| Workaround |
There is no workaround to prevent the incorrect success reporting. Users should manually review Protect dsmerror log files after each backup operation to verify that all expected files were successfully backed up, even when mmbackup reports success. Look specifically for ANS4047E errors and "Total number of objects failed" messages in the Protect client logs. |
|
6.0.1.1 |
mmbackup |
| IJ58974 |
High Importance
|
When multiple IOs with overlapping buffers are happening on same strip of a vtrack, the current code gets into a reconstruction deadlock due to a confusion in threads while deciding which buffers they should reconstruct & which buffers they should wait for others to reconstruct and complete the IO. We see long waiters saying "waiting for incompatible vdisk operations" in the dumps.
(show details)
| Symptom |
Hang/Deadlock/Unresponsiveness/Long Waiters |
| Environment |
ALL Operating System environments |
| Trigger |
While the IOs on same strip of same vtrack are in progress, a disk goes down triggering reconstruction in the IO threads. |
| Workaround |
NA |
|
6.0.1.1 |
ESS/GNR |
| IJ58894 |
High Importance
|
Log group resign happens when some error seen on the disk.
(show details)
| Symptom |
Abend/Crash |
| Environment |
Linux Only |
| Trigger |
Simply creating the new recovery group would have this problem and any medium writes with disk failures would end up causing the problem. |
| Workaround |
None |
|
6.0.1.1 |
GNR |
| IJ58983 |
High Importance
|
GPFS can halt due to logAssert(jniFlags & 0x02) or logAssert(id == myClusterId) in clusters where mmexpelnode has been invoked to expel other nodes and the persistent expel feature has not been disabled. This feature can be disabled with
(show details)
| Symptom |
Abend |
| Environment |
All |
| Trigger |
Numerous scenarios including unexpelling nodes or startup/restart/changes in existing nodes |
| Workaround |
The problem cannot occur in clusters where persistent expel feature is disabled. To disable thisfeature, run
mmchconfig "disablePersistExpelList=yes" -i
This must be run in both the local cluster (for any expelled nodes) and any storage clusters of which the expelled nodes were clients. |
|
6.0.1.1 |
GPFS core |
| IJ58839 |
High Importance
|
In large clusters with hundreds of nodes, network delays and inconsistencies in the delivery of QoS statistics and control messages can lead to race conditions. These race conditions may cause in-flight QoS configuration updates or refresh operations to hang while attempting to stop the existing QoS manager. Because the QoS shutdown thread holds filesystem-level mutexes, such a hang can cascade into cluster-wide waiters if other operations attempt to acquire the same filesystem-level locks.
(show details)
| Symptom |
Hang/Deadlock/Unresponsiveness/Long Waiters. |
| Environment |
ALL Operating System environments. |
| Trigger |
Frequent QoS configuration updates and refreshes in a large cluster with network delays, during heavy I/O workload. |
| Workaround |
None |
|
6.0.1.1 |
QoS |
| IJ59364 |
High Importance
|
mmfsckx reports false positive non-critical directory entry corruptions.
(show details)
| Symptom |
False positive output by mmfsckx |
| Environment |
All |
| Trigger |
Running mmfsckx on a busy system |
| Workaround |
Use offline fsck |
|
6.0.1.1 |
mmfsckx |
| IJ59361 |
Suggested |
A being deleted disk isn't removed from the performance disk address leading to FSErrBadDiskAddrIndex. That's because the performance pool isn't repaired when repairing the regular pool failed with E_NOREPLGRP when the file has two regular replicas.
(show details)
| Symptom |
Abend/Crash |
| Environment |
ALL Operating System environments |
| Trigger |
A file in a DAT file system has two regular replicas and the regular pool doesn't have enough failure groups. |
| Workaround |
Make sure there are enough failure groups for the reliable pool, and then run mmrestripefs or mmrestripefile. |
|
6.0.1.1 |
All Scale Users of a DAT file system |
| IJ58981 |
High Importance
|
Lookup of file is evicting file after data is appended to the uncached file.
(show details)
| Symptom |
Unexpected Results |
| Environment |
Linux Only |
| Trigger |
File data block in cache are zero after append operation is replicated to home. |
| Workaround |
None |
|
6.0.1.1 |
AFM |
| IJ59363 |
High Importance
|
GPFS registers a PR key on a disk that doesn't belong to GPFS any more
(show details)
| Symptom |
Unexpected Results/Behavior |
| Environment |
ALL Operating System environments |
| Trigger |
Because the underlying dev name can change on node reboot, CCR could get stale information and reserves key without checking the disks in CCR nsdmap have a valid nsdId. |
| Workaround |
No work around to prevent this problem from happening, one can remove the key after the fact. |
|
6.0.1.1 |
All Scale Users |
| IJ59463 |
HIPER |
Samba can deadlock if multiple SMB clients concurrently open the same files.
(show details)
| Symptom |
Hang/Deadlock/Unresponsiveness/Long Waiters |
| Environment |
ALL Linux OS environments |
| Trigger |
Open the same files multiple times through SMB clients connected to different nodes. The exact scenario is that two SMB clients connect to two different CES nodes and each open a different file. That grants oplocks for each open file.
If then both SMB clients open the file the other SMB client already has open, at the same time, then this results in two oplock break requests. If the oplock breaks overlap and each SMB client is not able to respond to the break request and both stay stuck in the open calls. A bug in the GPFS code prevents breaking out from this scenario. |
| Workaround |
The deadlock occurs due to the breaks happening for granted oplocks across multiple nodes. That provides multiple workarounds:
1) Not using oplocks at all avoid this problem at a possible cost of performance (mmsmb oplocks).
2) If data is only accessed through SMB, the cross-protocol integration can also be disabled (mmsmb gpfs:leases=no), but this is dangerous if the same data is also accessed outside of SMB.
3) Ensure that data for each SMB share is only accessed through one CES node removes the cross-node aspect. The simple approach would be moving all CES IP address to one node, but then that node needs to be able to handle the complete SMB workload. |
|
6.0.1.1 |
SMB |
| IJ58594 |
High Importance
|
Filesystem unmounts with error code 786 and message "encryptionHandled not set for inode" during I/O operations on encrypted filesystems accessed via remote mounts, resulting in filesystem unavailability and I/O failures.
(show details)
| Symptom |
Filesystem panic/unmount (error code 786) |
| Environment |
All (Linux, AIX, Windows)> |
| Trigger |
Encrypted filesystem with remote mount configuration during sustained I/O workload when token revoke operations occur (e.g., another node requesting access, periodic token management, or memory pressure triggering background sync/flush operations). |
| Workaround |
None |
|
6.0.1.0 |
File System Encryption |