Showing posts with label Red Alerts. Show all posts
Showing posts with label Red Alerts. Show all posts

Wednesday, November 23, 2016

Red Alert - IMS V12, V13 or V14 - potential for IMS to write incorrect log data

Yesterday, IBM issued a Red Alert on IMS. Here's the info. I'm just taking over the text from the Red Alert.

Title:

IMS V12, V13 or V14 - potential for IMS to write incorrect log data and/or over-write 64-bit (key 7) common storage which might belong to another address space.

Users Affected:

Users of IMS that are on V12 or later,
  • who are using log buffers in above-the-bar storage (BUFSTOR=64), or
  • have callers who supply an above-the-bar log record to the IMS Logger, such as:
  •       IMS 64-bit Fast Path buffer manager
          Installation-written or vendor-written software

Description:

The IMOVE macro (used by the IMS Logger) uses 32-bit instructions to advance the addresses of the source and target destinations. If the address being advanced is a 64-bit address which crosses a 4GB boundary, then the updated address is incorrect since only the low-half of the address is updated.

The two areas of IMS that are exposed to this are the IMS Logger and the 64-bit Fast Path buffer manager.

The IMS Logger is exposed to the problem if the storage allocated for the log buffers crosses a 4GB boundary. If this is the case, a log record will contain inaccurate data and an attempt will be made to inadvertently write to another location in memory. If this location is key 7, the data will be overwritten; if it is not, IMS will ABEND0C4.

The 64-bit Fast Path buffer manager is exposed to the problem if its BPND5 area crosses a 4GB boundary. To encounter the problem, not only would the area have to cross a 4GB boundary, but log record data must exist in the location where the boundary is crossed.

Recommended Actions:

(1) Determine exposure from IMS Logger or 64-bit Fast Path buffer manager
The IMS Support Center is shipping a diagnostic utility as a ++USERMOD which will determine the exposure of a given IMS system. This can help a customer gauge the urgency for which they should take action, if any. The utility runs as a stand-alone batch job and determines:
  1. If the log buffers are above the bar, and if the storage allocated for the log buffers crosses a 4GB boundary.
  2. If the 64-bit Fast Path buffer manager is enabled, and if the storage area which might contain a log record crosses a 4GB boundary
To download the utility, obtain it from:
Site: testcase.software.ibm.com
Directory: fromibm/im
File name: IM97624A.trs

The file contains ++USERMODs for IMS versions 12, 13, and 14. Instructions for running the utility are in a ++HOLD card along with its return codes.

If a customer cannot (or does not wish to) run the utility, instructions for determining the exposure via a series of /DIAGNOSE commands are in the ++HOLD card.

(2) Determine exposure from other programs which invoke the ILOG macro:
Installation-written or vendor-written programs can detect if they pass log records which reside above the bar by searching for the flag PRMLL64 which is set by the invoker of ILOG in the ILOG parameter list.

(3) Apply PTF if exposure is identified
If an exposure is identified, the exposed IMS should be:
  1. Brought down cleanly (for example, with the /CHE FREEZE or /CHE DUMPQ commands)

  2. Cold start IMS with the PTF applied for the appropriate version. The PTFs can be downloaded from ShopZ and are:
  3. V12: UI42725 (APRA PI71688)
    V13: UI42726 (APAR PI71701)
    V14: UI42685 (APAR PI71702)
    A cold start is necessary to prevent the reading of potentially inaccurate log data.

  4. After all previously exposed IMS systems in a data sharing group are successfully started with the PTF applied, if there is the possibility of needing to run a database recovery before regularly scheduled image copies will be taken, it is recommended that image copies be taken of all appropriate databases. This eliminates the possibility of reading potentially corrupt log data during a recovery.

If you haven't signed up to the Red Alerts by now, you really should do it. Just go over here.

Tuesday, March 15, 2016

Potential exposure for undetected loss of data for z/OS users writing data to zEDC compressed data sets using QSAM

Here's a new Red Alert. I'm taking over the content.

Abstract:
Potential exposure for undetected loss of data for z/OS users writing data to zEDC compressed data sets using QSAM on releases z/OS 1.13, 2.1 & 2.2.

Description
Data loss may occur when writing to a zEDC compressed data set using QSAM when CLOSE is issued and there exists a partial QSAM buffer that has not yet been written, and all allocated space in the data set is filled. In this case, the unwritten partial buffer may not be written after CLOSE processing obtains the new extent for the data set.

The problem does not apply to BSAM or other access methods.

Please see APAR OA50061 for additional information.

Recommended Actions 
++APAR for OA50061 should be applied for all environments with zEDC compressed data sets. This requires an IPL to activate.

You can find the APAR information over here where you can find some more info :
"During creation of a new extent during CLOSE processing, any
user blocks that have not been compressed and written out to the
zedc compressed dataset yet will not be written out.

This does not apply to BSAM processing since all outstanding
WRITE requests must be CHECK'd for completion before CLOSE is
issued."
For the moment it also says : "++APAR will be available soon".

If you haven't signed up to the Red Alerts by now, you really should do it. Just go over here.

Thursday, February 18, 2016

Two Flash Alerts for DS8870

Well, life goes on, and this week I saw two flash alerts passing by on DS8870. You can find them here and here. They're not that long, so I'm just taking over both of their contents.

DS8870 Global Mirror suspend caused by a Track Format Descriptor mismatch

Abstract
Global Mirror suspends caused by a microcode logic error introduced in R7.4 that results in a Track Format Descriptor mismatch. 

Content
Global Mirror suspends caused by a microcode logic error introduced in R7.4 that results in a Track Format Descriptor mismatch. Microcode is improperly setting a flag in a PPRC control block. This problem is pervasive in Global Mirror environments. R7.4 code levels below 87.41.44.0 and R7.5 levels below 87.51.23.4 are exposed to this issue.
Recommendation:
A mandatory ECA 712 is being released to the field, we recommend upgrade to code bundle R7.5 87.51.23.4

High Performance Flash Enclosure drive failures can cause loss of access

Abstract
IBM released microcode with improved error handling for DS8870 High Performance Flash Enclosure (HPFE) flash drive errors. 

Content
IBM has developed microcode enhancements for error handling in HPFE. DS8870 (R7.5 SP2.3 ) 87.51.23.4 microcode levels contain the changes to improve the reliability and availability of High Performance Flash Enclosures. This enhancement streamlines error detection and isolation when a failing Flash Drive exhibits excessive errors.

This change is designed to be concurrently installable on DS8870 presently running the R7.x families of microcode.

Recommendation:
A mandatory ECA 737 is being released to the field, we recommend upgrade to code bundle R7.5 SP2.3 87.51.23.4 as soon as possible.

ID: 312723 312369 313417


Wednesday, December 9, 2015

Flash : TS7700 potential data loss issue with back-end disk cache 3956-Cx9

Here's a flash about potential data loss on a TS7700 with back-end disk cache controller 3956-CS9/CC9 and disk cache expansion 3956-XS9/CX9.Here's part of its content :

Background:
Certain model disk drives within 3956-CS9/CC9 model controllers and 3956-XS9/CX9 model enclosures are exposed to a rare issue that can result in data loss. The problem is related to an internal integrity checking mechanism contained within certain model disk drives. There is a rare condition in which an underlying disk drive scanning routine can corrupt a disk sector some time after content has been previously written. (...)

Drive Model and Firmware Details:
For 3956-CS9/XS9 configurations, 3TB drives with model ST3000NM0043 are exposed to this issue when down level firmware is present. Firmware level EC5C or greater contains a fix for this issue.
For 3956-CC9/CX9 configurations, 600GB drives with model ST600MM026 are exposed to this issue when down level firmware is present. Firmware level E56F or greater contains a fix for this issue.

Solution (Procedure):
Vtd_exec.154 v1.41 or above should be installed on all affected systems.
Note: An updated version of vtd_exec.154 is expected to be released mid December. It is recommended to utilize the updated version if possible.

Alternatively upgrade the TS7700 to latest released code level R3.3 (8.33.0.45).

Either of these solutions must be applied by an IBM Service Representative.

Do check out the full information in the flash itself.

Tuesday, December 1, 2015

Flash - DS8000 with zHPF enabled and SyncSort MFX 2.1 failures

This is a flash for all current DS8000 systems ranging from DS8300 to DS8870. The DS8880 system only becomes available by the end of this week. You can find the flash over here. Below, I'm taking over the content of this flash.

Abstract

Running with zHPF enabled on a IBM System Storage DS8000®, can lead to host hangs, looping jobs and recursive errors on the IBM System Storage DS8000®system. 

Content

IBM System Storage DS8000® when running in a z/OS environment with zHPF enabled can result in:

1. In certain cases customers have experienced hosts to become unresponsive after sense is off-loaded from the DS8000 with zHPF enabled. This was seen only with SyncSort MFX 2.1. SyncSort is recommending APAR TY01136 for this issue. Until the APAR is applied, SyncSort is recommending the customer disable zHPF for SyncSort MFX 2.1 (use parameter LBPZHPF to bypass zHPF usage for an individual job and installation option SCOS=(204) to turn off all zHPF usage for SyncSort).
Licensed SyncSort MFX customers can reference Knowledge Base Article 46965 for more information about PTF TY01136.


2. In other situations customers have encountered repeating warmstarts, after software issues a zHPF Multi-Track Read command when the tracks are configured with more than 55 records per track, and there are multiple tracks to transfer. This results in the DS8000 incorrectly surfacing a Microcode Logic Error (MLE) which requires a warmstart to clear. This has been seen only with SyncSort MFX 2.1 software, although other software implementing similar chains would be exposed. After the MLE, software re-drives the command chain, resulting in repetitive MLEs that can result in system access loss. There is a DS8000 microcode fix in test scheduled for a near term release. In the interim, IBM is recommending that customers disable zHPF for SyncSort MFX 2.1 until the DS8000 fix is applied.

Examples of errors seen:

Message to Operator (WTO) notices WER529A and WER999A.
Missing Channel End and/or Missing Device End.
Host hangs/unresponsiveness.
DS8000 Microcode bundles containing the fix
DS8700 76.31.159.0
DS8800 86.31.184.0
DS8870 87.41.42.0
87.51.23.0

Tuesday, July 7, 2015

Red Alert for IBM DB2 V10 and DB2 V11 users

Here's a new Red Alert. I'm taking over part of the content.


Title:
For DB2 V10 and DB2 V11 users: DB2 log records may be written with invalid control information in certain rare cancel window situations and cause unpredictable results 

Abstract:
DB2 log records may be written with invalid control information in certain rare cancel window situations. This could cause log read errors, data replication issues, recovery issues, or even potential data loss. Both DB2 10 and 11 systems are affected, although the exposure is higher for DB2 11 systems that have not converted BSDS data sets to use 10 byte RBA values. DB2 9 and prior releases are not affected.

There's an entire description in the Red Alert itself. Let me just give you the recommended actions.

Recommended Actions for DB2 11 customers:
  1. Install APAR PI41630 (PTF UI28315). This fix will prevent log records with LRSN value of 'FFFFFFFFFBAD'x being created.
  2. Install APARFIX for PI42170 (AI42170) or PTF when available. This fix will prevent log records with invalid values being created, and it contains toleration code to recognize and recover from the incorrect values, including LRSN value of 'FFFFFFFFFBAD'x.
Recommended Actions for DB2 10 customers:
  1. Install APARFIX for PI42170 (AI42170) or PTF when available. This fix will prevent log records with invalid values being created, and it contains toleration code to recognize and recover from the incorrect values, including LRSN value of 'FFFFFFFFFBAD'x if the system fallback from DB2 11.
 If you haven't signed up to the Red Alerts by now, you really should do it. Just go over here.

Thursday, February 12, 2015

Security Bulletin: GNU C library (glibc) vulnerability affects DS8000 and XIV

IBM issued some security bulletins referring to all models of the DS8000 and the Gen2 and Gen3 of the XIV.

In general, here's what it's about :

"Summary

GNU C library (glibc) vulnerability that has been referred to as GHOST affects DS8000

Description: 

The gethostbyname functions of the GNU C Library (glibc) are vulnerable to a buffer overflow. By sending a specially crafted, but valid hostname argument, a remote attacker could overflow a buffer and execute arbitrary code on the system with the privileges of the targeted process or cause the process to crash. The impact of an attack depends on the implementation details of the targeted application or operating system. This issue is being referred to as the "Ghost" vulnerability."

For more information and fixes, please refer to the appropriate flashes.
For DS8000 : link.
For XIV Gen2 : link.
For XIV Gen3 : link.

Monday, August 18, 2014

Red Alert - z/OS 2.1 DFSORT records out of sequence

I know I'm a bit late with this one but I still want to mention it, just in case you might've missed it.

Red Alert : z/OS 2.1 DFSORT records out of sequence

Abstract:

There is a potential exposure for out of sequence records with DFSORT for users on release z/OS 2.1.

Description:

At z/OS 2.1 code levels, DFSORT is intermittently returning records out of sequence. There is no data loss, but records may be returned out of sequence to the DFSORT output file. If VERIFY=YES is set at the installation level, out of sequence conditions are already being detected. This problem only occurs in z/OS 2.1. No prior releases of DFSORT are affected.

Please see APAR PI22817 for more details and latest information.

Users affected :

All z/OS 2.1 DFSORT (HSM1L00) users who SORT data with DFSORT using the performance path may be affected if there is insufficient virtual storage below the line at the time of execution. In other words the potential for error exists for all users of DFSORT SORT function on z/OS 2.1, but not all users will experience the problem.

Recommended Actions:

Enable VERIFY=YES at the installation level to detect out of sequence conditions. Affected jobs can be rerun with DEBUG $NOPFP$ to circumvent the issue.

In addition, a ++APAR is available to disable the affected performance path, as a temporary circumvention, using a new DFSORT installation option.

See APAR PI22817 for details. 


If you haven't signed up to the Red Alerts by now, you really should do it. Just go over here.

Monday, October 28, 2013

Red Alert : Unable to partition systems from the sysplex on z/OS 1.12 with PTF UA69636 (APAR OA41661) applied

Here's a new Red Alert. I'm just taking over the content.

Unable to partition systems from the sysplex on z/OS 1.12 with PTF UA69636 (APAR OA41661) applied

Abstract:

Unable to partition systems from the sysplex on z/OS 1.12 with PTF UA69636 (APAR OA41661) applied

Description:

When attempting to partition a system (system 1) from the sysplex, the partitioning process on the monitor system which is managing the removal (system 2) may hang.
System 1 cannot be removed from the sysplex. System 2 remains functional, but cannot continue with the removal of system 1.

Affected Environments:

Monitor system (system 2): z/OS running at z/OS V1R12 (HBB7770) with PTF UA69636 (APAR OA41661) applied.
Outgoing system (system 1): Any supported z/OS level with the SYSSTATDETECT function enabled.
Please see APAR OA43435 for more details.

Recommended Actions:

  1. Disable the SYSTATDETECT function on all systems in the sysplex via SETXCF FUNCTIONS,DISABLE=SYSSTATDETECT
  2. Please retrieve PTF package - PTF.UA71120 from: testcase.boulder.ibm.com under fromibm/mvs with FB, 80 3200 in binary. Apply this PTF to all applicable systems.
  3. The SYSSTATDETECT function should remain disabled on all systems until all affected systems are running with UA71120.
  4. If the reported problem has already occurred, contact Level 2 to determine best action.

If you haven't signed up to the Red Alerts by now, you really should do it. Just go over here.

Monday, June 3, 2013

Updated alerts for Red Alert APAR OA42277

Last week I posted on a 'Red Alert : Potential loss of data for DB2 v10 users on z/OS 1.12 and 1.13'. Now there are two updates on this Alert.
You can check them out over here :
I'm taking over the content of this last Red Alert.

Abstract:

An error has been discovered in a Pre Req for Red Alert APAR OA42277

Potential loss of access to data in DB2 or MQSeries Log data sets greater than 2gb on z/OS 1.12, 1.13 after application of PTF's UA68675 & UA68676 for APAR OA39870 and base z/OS 2.1.

Description:

Any application using the RBA interface to Media Manager accessing Extended Format (EF) VSAM data sets greater than 2gb may lose access to the data past 2gb. Applications using the relative CI interface are not affected. Data sets defined with Extended Addressability are also not affected. This includes all current versions of DB2 and MQSeries and potentially other users of Media Manager's RBA Interface to access EF VSAM Data Sets greater than 2gb, such as IMS FastPath DEDB and other utility programs which may access these data sets.

A usermod/method is currently being developed to be able to access the data above 2gb.

Please see APAR OA42425 for more details.

Recommended Actions:

  1. Do not install PTF's for Red Alert APAR OA42277 or PTF's for Pre Req OA39870.
  2. To avoid the problem in Red Alert APAR OA42277, zDMF users should refrain from using zDMF until complete PTF chain can be installed. The risk of OA42277 is extremely small for those customers without zDMF.
  3. If PTF's for OA39870 are installed on z/OS 1.12 or 1.13 please remove as soon as possible.
  4. Extended format VSAM datasets greater than 2gb, such as DB2 Logs, MQ Logs and IMS FastPath DEDB, accessed with Media Manager using the RBA interface while the fix for OA39870 was installed, should be verified. Please see APAR OA42425 for instructions.
  5. Monitor APAR OA42425 for updates regarding the availability of the usermod and recovery procedures.
  6. z/OS 2.1 users will need to install the ++APAR for OA42425 when available.

Thursday, May 30, 2013

Red Alert : Potential loss of data for DB2 v10 users on z/OS 1.12 and 1.13

Here's a new Red Alert. I'm just taking over the content.

Potential loss of data for DB2 v10 users on z/OS 1.12 and 1.13 releases

Users affected:
DB2 V10 and PDSE users potentially affected by error in z/OS Media Manager IO recovery. Users of zDMF running zHPF may be particularly exposed because all I/O requests are intercepted by zDMF and returned to Media Manager in error for IO redrive using FICON, however any I/O errors, including interface control check and unit checks, that are retried by the media manager are also subject to the problem.

Description:
There are certain situations where the Media Manager will redrive the failed channel program with a different type of channel program. In some highly timing dependant cases, such as with I/Os that cross extents, the redrive may not occur correctly, and no error is presented to the caller, potentially leading to data being down level on disk. The problem may occur for DB2 V10 or PDSE data sets. Please see APAR OA42277 for more details.

Recommended Actions:
Apply ++APAR or PTFs for OA42277. Activation requires re-IPL of the system.

Also please do not use zDMF until the fix for OA42277 is applied.

If you haven't signed up to the Red Alerts by now, you really should do it. Just go over here.

Friday, April 5, 2013

Flash Alert : TS7700 Potential Performance Issue with R3.0 code and 3957 VEA/V06 models

For those who are in this situation, you might have a look at this before upgrading to the R3.0 (8.30.x.x) release level. I'm taking over the info from the alert you can also find over here.

Abstract

TS7700 8.30.x.x levels may cause a performance degradation when installed on the TS7720 (3957-VEA) or the TS7740 (3957-V06) hardware model base.

Content

MachineType/Model affected: TS7720 (3957-VEA), TS7740 (3957-V06).

It has been determined by IBM that TS7700 3.0 Release levels (8.30.x.x) may cause a performance degradation when installed on the TS7720 (3957-VEA) and the TS7740 (3957-V06) hardware model base. Customers considering an upgrade to the R3.0 (8.30.x.x) release level on these models, or installing a TS7720 expansion frame, 3952-F05 w/ FC 7332, are strongly advised to have a SCORE request submitted, allowing IBM to assess current performance requirements and to provide a recommended upgrade path. Environments that run performance-sensitive, or high-throughput workloads on VEA or V06 hardware platforms will require special consideration. Customers should contact their IBM sales team for assistance submitting the SCORE request so that the appropriate technical reviews can be completed.

This does not apply to customers with TS7720 (3957-VEB) and/or TS7740 (3957-V07), who may upgrade to R3.0 without restriction.

Tuesday, March 26, 2013

Red Alert : PTF UK91435 required for DB2 10 for z/OS in a Data Sharing environment

Here's a new Red Alert. I'm just taking over the content.

PTF UK91435 (APAR PM79520) is required for DB2 10 for z/OS customers in a Data Sharing environment

Description:
The subject PTF addresses a potential data loss in a DB2 10 Data Sharing environment.

The problem is related to lost spacemap updates in Data Sharing. Data modifications (insertions, mass deletions) may have been lost, and if so, will not be recoverable by standard recovery procedures. The modifications will remain recorded on the DB2 recovery log and can be recovered using tools such as the Log Analysis Tool. Rebuilding the index will make the Index consistent with the data but will not recover data that is missing nor attend to erroneously present data. The same is true of reorganizing the object. If in the interest of system availability indexes are rebuilt or the data is reorganized prior to determining if there is data loss, it is recommended that this action be followed with analysis and possible remedial action as outlined below.

Possible symptoms include: Incorrect output; ABEND04E RC00C90101, RC00C90102, RC00C90105, or RC00C902xx in various CSECTs; data/index inconsistencies reported by the CHECK INDEX utility; and, page regression reported by the DSN1LOGP utility.

Recommended Actions:
IBM recommends that all DB2 10 for z/OS Data Sharing customers apply PTF UK91435 as soon as possible.

Customers may validate their data by using CHECK INDEX or CHECK DATA utilities. Any issues uncovered should be reported to IBM Software Support to determine root cause. The data may not be recoverable by standard recovery procedures.

Refer to the cover letter for UK91435 for further details regarding this issue.

Tuesday, March 12, 2013

DS8100-DS8300-DS8700 - DDM Firmware Issue – Possible Undetected Data Loss or Data Error

Here's a flash alert which is issued for three DS8000 models : DS8100, DS8300 and DS8700. You can find it over here.

I'm taking over the most important parts. Go to the alert itself for more details.

Abstract

Certain disk drive modules (“DDMs”) shipped between April 2010 and January 2013, running DDM firmware levels F520, F522, or F527, may be exposed to a possible undetected data loss or data error during a proximal write. (The “proximal write” feature does a skip operation on the data transfer from DRAM to disk to improve performance.) This issue occurs when the starting logical Block Address (“LBA”) is a reassigned LBA. A firmware update designed to address this issue is now available.
Note: DS8800s and DS8870s are not exposed to this issue. DS8000 DDMs that use drive-level encryption are also not exposed to this issue.

Content

Fix / Mitigation Options
A Concurrent DDM Firmware update with firmware F529 using Install Corrective Service (ICS) CD for machines running Bundles 64.20.xx.xx or higher (8100/8300) and 76.20.xx.xx or higher (8700) is now available. Clients with DS8000s below these minimum bundles and deciding to update the DDM firmware need to either perform a code load to one of the bundles identified above and then apply ECA 866, or contact IBM to evaluate other options.

Identification Methods
  1. DDM firmware levels can be queried by an IBM Service Support Representative (“SSR”) using the service panels, or by clients using CLI commands (see examples below).
  2. An Info Alert is being released to notify SSRs of subsystems containing DDMs with F520, F522, or F527 firmware. An as required ECA [“Engineering Change Action"] 866 is also being released to provide SSRs with instructions describing how to update the DDM firmware for clients that request this update
  3. If clients determine they have a system with the affected DDMs, IBM Service can be contacted to schedule the update.
  4. Please contact IBM Service, or contact the DS8000 Quality team at the following Email address, DS8KQWT@US.IBM.COM or DS8000QWTeam/Tucson/IBM, for any questions or to request additional assistance.

Monday, October 8, 2012

Red Alert : Unpredictable task failures and/or system outages with OMEGAMON XE on z/OS

Here's a new Red Alert:

Unpredictable task failures and/or system outages with OMEGAMON XE on z/OS

Description:
There is an exposure for storage overlays in common and private storage that can cause unpredictable task failures and/or system outages. All users of the OMEGAMON XE on z/OS product versions V420 or V510 with the following PTFs for OA39579 applied are affected:
  • HKM5420 - UA66217
  • HKM5510 - UA66218
The PTFs for OA39579 introduced the two problems described below. The PTFs for OA40262 resolved only the first problem. Fixes for OA40497 are also required to address the complete problem.
  1. Random storage overlays due to an incorrect register used in an instruction. One symptom of this problem is many LOGREC entries for ABEND0C4 in module KXDWLCON. The overlay contains the character string 'AIO'.
  2. An incomplete modification to an ENF exit causes ECSA storage overlays when a WLM policy switch is done.
Recommended Actions:
Apply the appropriate PTFs for OA40497 which will ensure both problems are resolved:
  • HKM5420 - UA66779
  • HKM5510 - UA66780

If you want to have an overview of all past Red Alerts, then take a look over here. You can also subscribe on that same page so you'll be notified of any future Red Alert.

Wednesday, September 5, 2012

Red Alert : potential exposure for loss of data or data set corruption for DFSMS VSAM/RLS users on z/OS 1.13

Well, life goes on, even after the announcement of a new system. Here's a new Red Alert:

Potential exposure for undetected loss of data or data set corruption for DFSMS VSAM/RLS users on release z/OS 1.13

Users affected:
All zOS 1.13 (HDZ1D10) VSAM/RLS users who are sequentially erasing records from a KSDS while simultaneously updating records via a concurrent request.

Products affected:
zOS 1.13 DFSMS VSAM/RLS (HDZ1D10)

Description:
It is possible for an inserted or updated record to be unintentionally erased or down leveled if concurrent RLS access of a KSDS has sequential erase activity and inserts/updates taking place at the same time. In order for data to be lost, at least one sequential erase must remove all records in a CI, such that CI reclaim is initiated for that CI. Subsequent sequential requests may then overlay other updates to the data set, causing the previously updated records to be lost. The job that updated the record may not receive an error indication, or may receive unexpected logic errors such as record not found. EXAMINE may also report key sequence errors.

Please see APAR OA40253 for additional information and actions to determine exposure.

Recommended Actions:
Apply ++APAR for OA40253.

If you want to have an overview of all past Red Alerts, then take a look over here. You can also subscribe on that same page so you'll be notified of any future Red Alert.

Thursday, July 19, 2012

Possible DS8700/DS8800 abort condition during zOS writes to a FlashCopy target when experiencing uncorrectable SAN link errors

I'm just reproducing this Flash (alert) which you can find over here.

Abstract
In certain circumstances, CKD (zOS) host writes to a FlashCopy Target in conjunction with uncorrectable SAN link errors can result in repetitive microcode logic errors and possible loss of access to the DS8700/DS8800.

Content
The issue specifically requires a combination of CKD writes to a FlashCopy Target volume in addition to uncorrectable SAN link errors. Most notable use for a zOS

Host to write to a FlashCopy Target is during disaster recovery testing in a Global Mirror (GM) or Metro Global Mirror (MGM) environment. Typically, for disaster
recovery tests, at the Global Mirror secondary, a practice copy will be created from the consistent copy of the data. This practice copy is then made available for
Host I/O directly to it.

Another common use for a FlashCopy target, tape backup, is not affected by this issue as tape backup only does Host reads and not writes.

Release 6.1, Release 6.2 (prior to Service Pack 2.1), Release 6.3 (prior to Service Pack 1) on the DS8700 and DS8800 platforms are affected.

Mitigation:
Two aspects of mitigation for this problem:

1. Avoid any activity which involves CKD I/O to a FlashCopy Target. Avoid disaster recovery testing in Global Mirror environments which involve using a practice
copy from the Global Mirror target until fix is delivered.

2. In the case of an actual disaster where the Global Mirror primary site is no longer available, the user can invoke a fast reverse restore from the consistent copy
(D volume) of the data back to the C volume and run the zOS host specifically to the C volume. The C volume in this example is not a FlashCopy target, it is a
regular volume.

Note: If the application waits for the background copy to complete prior to initiating other CKD writes to the FlashCopy Target volume, there is no risk of the abort mentioned above.

Resolution:
A fix has been available in microcode bundles 76.20.94.0 or higher for DS8700s and in 86.20.117.0 or higher for DS8800s. A fix is also now available in release 6.3 microcode bundles 76.31.17.0 or higher for DS8700s and in 86.31.26.0 or higher for DS8800s.

Monday, July 2, 2012

Red Alert : JES2 Potential Loss of Spool data on z/OS 1.11 and 1.12

Here's a new Red Alert:

JES2 Potential Loss of Spool data on z/OS 1.11 and 1.12

Description:
The fix for APAR OA36256 (RSU1112 PTFs UA61942, UA61943 UA61944 on HJE7760, HJE7770, and HJE7780 respectively) widened a timing window during JES2 initialization processing such that an initializing JES2 member may not obtain the correct status of other multi-access spool (MAS) members. As a result, this system's view of a spool volume may differ from the rest of the MAS. Consequently, later HALTING or DRAINING actions against the spool volume may result in incomplete cleanup.

PE APAR OA39737 will address the timing window and ensure the initializing member has the most accurate status of the MAS during initialization. In addition, APAR OA38016 will address spool errors caused by the timing window that may result in potential loss of spool data. These types of spool errors are already corrected in z/OS 1.13.

Please see OA39737 and OA38016 for more details or updates.

Recommended Actions:
  • If PTF for OA36256 is applied, please avoid putting a spool volume into DRAINING or HALTING state. New volumes can be added or started without exposure.
  • If a Spool Drain or Halt must be done, Level 2 can check dumps of JES2 to determine if the spool volume is exposed.
  • If OA36256 is applied and a spool volume is already in DRAINING or HALTING state, please remove OA36256 and then (rolling) WARM start each JES2 member.

If you want to have an overview of all past Red Alerts, then take a look over here. You can also subscribe on that same page so you'll be notified of any future Red Alert.

Monday, June 4, 2012

Potential DS8000 loss of access with VMWare 4.1 and higher with VAAI Zero Blocks/Write Same enabled

Perhaps a little bit out of my league but nevertheless worth mentioning :

Abstract

The extended features of VAAI ( Zero Blocks/Write Same ) are not supported by DS8000 at this time, but are enabled by default on ESX 4.1 or higher.

Content

IBM has received reports of issues when running VAAI Zero Blocks/Write Same commands with DS8000 that have caused application problems including a loss of access to data. Clients are advised to disable VAAI's use of the Zero Blocks/Write Same command to avoid the possibility of an issue, and the procedure to do so is described in the Mitigation information below.

Mitigation

To avoid any issues, use the following to disable the Zero Blocks/Write Same function of VAAI (HardwareAcceleratedInit parameter). Other functions of VAAI may remain enabled for use on other storage systems. All OTHER non VAAI features of VMWare ESX 4.1 or higher are certified with DS8000 (all models) and function normally.

Here's a KB for disabling the Zero Blocks/Write Same command. - http://kb.vmware.com/selfservice/microsites/search.do?language=en_US&cmd=displayKC&externalId=1033665

Resolution/Support

IBM is working to address the issues of running VAAI Zero Blocks/Write Same functions with DS8000, and will send an update out when this has been done.

Friday, March 16, 2012

Red Alert : Potential data loss when executing data recovery in a DB2 10 (NFM) Data Sharing environment

Here's a new Red Alert:

PTF UK76697 (APAR PM51093) is required for DB2 10 for z/OS NFM (New Function Mode) customers executing data recovery in a Data Sharing environment

Description:

IBM has become aware of potential data loss when executing data recovery in a DB2 10 Data Sharing environment.

The problem is related to Fast Log Apply processing when LRSN values are used for log record sequencing and may result in some log records not being processed when duplicate LRSN values are encountered. This situation applies to the RECOVER and RESTORE SYSTEM utilities and LPL and GRECP recovery. The REORG utility and DB2 restart processing are not exposed to this issue.

Possible symptoms include ABEND04E RC00C90102 DSNIBHRE:0C22 accompanied by MSGDSNI012I PAGE LOGICALLY BROKEN or other consistency errors during log apply. Other abends or errors may be encountered following log apply processing, including errors reported by CHECK INDEX or CHECK DATA.

Customers may validate their data by using CHECK INDEX or CHECK DATA utilities and any issues should be reported to IBM Software Support to determine root cause.

Recommended Actions:

It is recommended that all DB2 10 for z/OS data sharing customers apply PTF UK76697 as soon as possible. Customers planning to migrate to DB2 10 NFM should ensure that UK76697 is applied before they do so. The corrective maintenance can be applied incrementally across members of a data sharing group. A group wide restart is not required.

Fast Log Apply processing may be disabled to prevent further exposure to this problem before UK76697 can be applied. The process for doing so is as follows, and should be repeated for each subsystem:

  1. Edit macro DSN6SPRC in the SDSNMACS library.
  2. Change &SPRMFLB SETC '10' to &SPRMFLB SETC '0'.
  3. Run job DSNTIJUZ to reassemble and relink the new zparm.
  4. Recycle DB2 to load the new zparm.

Refer to the cover letter for UK76697 for further details regarding this issue.


If you want to have an overview of all past Red Alerts, then take a look over here. You can also subscribe on that same page so you'll be notified of any future Red Alert.