Showing posts with label vmware. Show all posts
Showing posts with label vmware. Show all posts

Wednesday, December 11, 2024

VMware Health Analyzer - how to download and register the tool

Are you looking for VMware Health Analyzer? It is not easy to find it so here are links to download and register the tool to get the license.

Full VHA download: https://docs.broadcom.com/docs/VHA-FULL-OVF10

Collector VHA download: https://docs.broadcom.com/docs/VHA-COLLECTOR-OVF10

Full VHA license Register Tool: https://pstoolhub.broadcom.com/

I publish it mainly for my own reference but I hope other VMware community folks find it useful.

Monday, August 27, 2018

VMworld 2018 announcements

In this post, I would like to summarize the coolest VMworld 2018 announcements.

Project Dimension
On-premise managed vSphere infrastructure in a cloudy fashion. Project Dimension will extend VMware Cloud to deliver SDDC infrastructure and hardware as-a-service to on-premises locations.  Because this will be a service, it means that VMware can take care of managing the infrastructure, troubleshooting issues, and performing patching and maintenance. For more info read this blog post.

Project Magna
Project Magna will make possible a self-driving data center based on machine learning. It is focused on applying reinforcement learning to a data center environment to drive greater performance and efficiencies. The demonstration illustrated how Project Magna can learn and understand application behavior to the point that it can model, test, and then reconfigure the network to a make it more optimal to improve performance. Project Magna relies on artificial intelligence algorithms to help connect the dots across huge data sets and gain deep insights across applications and the stack from application code, to software to hardware infrastructure, to the public cloud and the edge.

vSphere Platinum
VMware vSphere Platinum is a new edition of vSphere that delivers advanced security capabilities fully integrated into the hypervisor. This new release combines the industry-leading capabilities of vSphere with VMware AppDefense, delivering purpose-built VMs to secure applications. For more info read this blog post.

VMware ESXi 64-bit Arm Support. 
ESXi will probably run on Cavium ThunderX2 servers. Cavium ThunderX2 has very interesting specifications.

vSphere 6.7 Update 1
VMware announced vSphere 6.7 Update 1, which includes some key new and enhanced capabilities. Here are some highlights:
  • Fully Featured HTML5-based vSphere Client
  • Enhanced support for NVIDIA Quadro vDWS powered VMs (vSphere vMotion with NVIDIA Quadro vDWS vGPU powered VMs)
  • Support for Intel FPGA 
  • New vCenter Server Convergence Tool (allows migration from an external PSC architecture into embedded PSC architecture and also combine, merge, or separate vSphere SSO Domains)
  • Enhancements for HCI and vSAN
  • Enhanced vSphere Content Library (import of OVA templates from a HTTPS endpoint and local storage, native support of VM templates)
For more info read this blog post.

VMware vSAN 6.7 Update 1
vSAN 6.7 U1 will be available together with vSphere 6.7 U1. Here are some highlights:
  • Firmware Updates through VUM
  • Cluster Quickstart wizard
  • UNMAP support (capability of unmapping blocks when the Guest OS sends an unmap/trim command)
  • Mixed MTU support for vSAN Stretched Clusters (different MTU for Witness traffic then vSAN traffic)
  • Historical capacity reporting
Amazon Relational Database Service (RDS) on VMware
AWS and VMware Announce Amazon Relational Database Service on VMware. It is a database as a service managed by Amazon. Amazon RDS on VMware will be generally available soon and will support Microsoft SQL Server, Oracle, PostgreSQL, MySQL, and MariaDB databases. Read announcement or register for preview here.

VVols support for SRM is now officially on the roadmap
It is not coming in the latest SRM version but it is officially in the roadmap so VMware announced the commitment to develop it soon. [Source]

VMware vCloud Director 9.5
The new vCloud Director 9.5 enhances easy and intuitive cloud provisioning and consumption by adding highly-requested capabilities including self-service data protection, disaster recovery, and container-orchestration for cloud consumers, along with multi-site management, multi-tenancy and cross-platform networking for Cloud Providers.  6 Key New Innovations in vCloud Director 9.5:
  • Cross-site networking improvements powered by deeper integration with NSX
  • Initial integration with NSX-T
  • Additional integration with NSX including e.g. the possibility to stretch networks across virtual datacenters on different vCenter Servers or vCD instances residing at different sites right from the UI.
  • Cross-platform networking for Cloud Providers. Makes it possible for NSX-T and NSX-V managers in same vCD instance to create isolated logical L2 networks with a directly connected network.
  • Full transition to an HTML5 UI for the cloud consumer
  • Improvements to role-based access control
  • Natively integrated data protection capabilities, powered by Dell-EMC Avamar
  • vCD virtual appliance deployment model
  • Container-orchestration for cloud consumers. Deploy both VMs and containers, consumed via Kubernetes.
  • Data protection capabilities. EMC Avamar is added to the vCD UI to make it easier for end consumers to manage these tasks. This is made possible based on the extensible tools available via vCD which means you (software vendor, cloud provider) can publish services to the vCD UI as needed. Hopefully, other vendors will follow.
For more information read VMware blog.

Introducing VMware Cloud Provider Pod: Custom Designs
VMware announced a new product that will revolutionize the deployment of Cloud Provider environments through the first flexible, validated VMware cloud stack with 1-click deployment: VMware Cloud Provider Pod. Cloud Provider environments are complex to deploy thanks to interoperability, scalability, reliability and performance issues that constantly plague cloud admins and architects. It is a time-consuming process that takes weeks to months. Yes, there are “one-click” deployment products out there, but these are rigid, have stringent hardware compatibility requirements, and end up creating yet another datacenter silo to manage. The Cloud Provider Pod has been designed to deliver three key capabilities:

  • Allows Cloud Providers to design a custom cloud environment of their choice
  • Automates the deployment of the designed cloud environment in adherence with VMware Validated Designs for Cloud Providers
  • Generates customized documentation and guidelines for their environment that radically simplifies operations
For more information read VMware blog.

VMworld US 2018 General Sessions & Breakout Sessions Playback
You can watch general sessions and also technical (breakout) session online.  General sessions on VMworld US 2018 is available at https://www.vmworld.com/en/us/learning/general-sessions.html

A nice summary list of all VMworld US 2018 technical (breakout) session with the respective video playback & download URLs is available at https://github.com/lamw/vmworld2018-session-urls/blob/master/vmworld-us-playback-urls.md

Wednesday, April 18, 2018

What's new in vSphere 6.7

VMware vSphere 6.7 has been released and all famous VMware bloggers released their blog posts about new features and capabilities. It is worth to read all of these blog posts as each blogger is focused on a different area of SDDC so it can give you a broader context to newly available product features and capabilities. Anyway, industry veterans should start reading product Release Notes and official VMware blog posts first.

Please note, that this blog post is just an aggregation of information published in other places. All used sources are listed below.

Release Notes:
vSphere 6.7 Release Notes

VMware KB:
Important information before upgrading to vSphere 6.7 

VMware official blog posts:
Introducing VMware vSphere 6.7!
Introducing vCenter Server 6.7
Introducing Faster Lifecycle Management Operations in VMware vSphere 6.7
Introducing vSphere 6.7 Security
What’s new with vSphere 6.7 Core Storage
vSphere 6.7 Videos

Community blog posts:
Emad Younis : vCenter Server 6.7 What’s New Rundown
Duncan Epping : vSphere 6.7 announced!
Cormac Hogan : What's new in vSphere and vSAN 6.7 release?
Cody Hosterman : What's new in core storage in vSphere 6.7 part I: in-guest unmap and snapshots
Cody Hosterman :  What's new in core storage in vSphere 6.7 part V: Rate control for automatic VMFS unmap
William Lam : All vSphere 6.7 release notes & download links
Florian Greh (Virten) : VMware vSphere 6.7 introduces Skylake EVC Mode
Florian Greh (Virten) : New ESXCLI Commands in vSphere 6.7

So after reading all resources above let's aggregate and document interesting features area by area.

vSphere Management

vCenter with embedded platform services controller in enhanced linked mode. This is nice because you can leverage "vCenter Server High Availability" to achieve higher availability for PSC without the external load balancer. All benefits listed below.
  • No load balancer required for high availability and fully supports native vCenter Server High Availability.
  • SSO Site boundary removal provides flexibility of placement.
  • Supports vSphere scale maximums.
  • Allows for 15 deployments in a vSphere Single Sign-On Domain.
  • Reduces the number of nodes to manage and maintain.

vSphere 6.7 introduces vCenter Server Hybrid Linked Mode, which makes it easy and simple for customers to have unified visibility and manageability across an on-premises vSphere environment running on one version and a vSphere-based public cloud environment, such as VMware Cloud on AWS, running on a different version of vSphere.

vSphere 6.7 also introduces Cross-Cloud Cold and Hot Migration, further enhancing the ease of management across and enabling a seamless and non-disruptive hybrid cloud experience for customers.

vSphere 6.7 enables customers to use different vCenter versions while allowing cross-vCenter, mixed-version provisioning operations (vMotion, Full Clone and cold migrate) to continue seamlessly.

vCenter Server Appliance (VCSA) Syslog now supports up to three syslog forwarding targets.

The HTML5-based vSphere Client provides a modern user interface experience that is both responsive and easy to use and includes 95% of functionality available in Flash Client. Some of the newer workflows in the updated vSphere Client release include:
  • vSphere Update Manager
  • Content Library
  • vSAN
  • Storage Policies
  • Host Profiles
  • vDS Topology Diagram
  • Licensing
PSC/SSO CLI (cmsso-util) has some improvements. Repointing an external vCenter Server Appliance across SSO Sites within a vSphere SSO domain is supported. Repoint of vCenter Server Appliance across vSphere SSO domains is also supported. This is huge! It seems that SSO domain consolidation is now possible. The domain repoint feature only supports external deployments running vSphere 6.7. The repoint tool can migrate licenses, tags, categories, and permissions from one vSphere SSO Domain to another.

Brand-new Update Manager interface that is part of the HTML5 Web Client. The new UI provides a much more streamlined remediation process. 


New vROps plugin for the vSphere Client. This plugin is available out-of-the-box and provides some great new functionality. When interacting with this plugin, you will be greeted with 6 vRealize Operations Manager (vROps) dashboards directly in the vSphere client! 


Compute

vSphere 6.7 delivers a new capability that is key for the hybrid cloud, called Per-VM EVC. Per-VM EVC enables the EVC (Enhanced vMotion Compatibility) mode to become an attribute of the VM rather than the specific processor generation it happens to be booted on in the cluster. This allows for seamless migration across different CPUs by persisting the EVC mode per-VM during migrations across clusters and during power cycles.

A new EVC mode (Intel Skylake Generation) has been introduced.  Compared to Intel "Broadwell " EVC mode, the Skylake EVC mode exposes following additional CPU features:
  • Advanced Vector Extensions 512
  • Persistent Memory Support Instructions
  • Protection Key Rights
  • Save Processor Extended States with Compaction
  • Save Processor Extended States Supervisor

Single Reboot when updating ESXi hosts. It is reducing maintenance time by eliminating one of two reboots normally required for major version upgrades.

vSphere Quick Boot is a new innovation that restarts the ESXi hypervisor without rebooting the physical host, skipping time-consuming hardware initialization (aka POST, Power-On Self Tests).

New ESXCLI Commands. In vSphere 6.7 the command line interface esxcli has been extended with new features. vSphere 6.7 introduced 62 new ESXCLI commands including:
  • 3 Device
  • 6 Hardware
  • 1 iSCSI
  • 14 Network
  • 14 NVMe
  • 2 RDMA
  • 9 Storage
  • 6 System
  • 7 vSAN
for more information look here.

Fault Tolerance maximums increased. Up to 8 Virtual CPUs per virtual machine and up to 128 vRAM per FT VM. For more info look at https://configmax.vmware.com/ (ESXi Host Maximums)

Storage

Support for 4K native HDD. Customers may now deploy ESXi on servers with 4Kn HDDs used for local storage (SSD and NVMe drives are currently not supported). ESXi providing a software read-modify-write layer within the storage stack allowing the emulation of 512B sector drives. ESXi continues to expose 512B sector VMDKs to the guest OS. Servers having UEFI BIOS can boot from 4Kn drives.

XCOPY enhancement. XCOPY is used to offload storage-intensive operations such as copying, cloning, and zeroing to the storage array instead of the ESXi host. With the release of vSphere 6.7, XCOPY will now work with specific vendor VAAI primitives and any vendor supporting the SCSI T10 standard. Additionally, XCOPY segments and transfer sizes are now configurable. By default, the Maximum Transfer Size of an XCOPY ranges between 4MB-16MB. In vSphere 6.7, through the use of PSA claim-rules, this functionality is extended to additional storage arrays. Further details should be documented by particular storage vendor.

Configurable Automatic UNMAP. Automatic UNMAP was released with vSphere 6.5 with a selectable priority of none or low. Storage vendors and customers have requested higher, configurable rates rather than a fixed 25MBps. With vSphere 6.7 we’ve added a new method, “fixed” which allows you to configure an automatic UNMAP rate between 100MBps and 2000MBps, configurable both in the UI and CLI. I recommend reading this blog post for details how it works on Pure Storage.

UNMAP for SESparse. SESparse is a sparse virtual disk format used for snapshots in vSphere as a default for VMFS-6. In this release, automatic space reclamation for VM’s with SESparse snapshots on VMFS-6 is provided. This only works when the VM is powered on and only affect the top-most snapshot.

VVols enhancements. As VMware continues the development of Virtual Volumes, in this release is added support for IPv6 and SCSI-3 persistent reservations. With end-to-end support of IPv6, this enables organizations, including government, to implement VVols using IPv6. With SCSI-3 reservations, this substantial feature allows shared disks/volumes between virtual machines across nodes/hosts. Often used for Microsoft WSFC clusters, with this new enhancement it allows for the removal of RDMs!

Increased maximum number of LUNs/Paths (1K/4K LUN/Path). The maximum number of LUNs per host is now 1024 instead of 512 and the maximum number of paths per host is 4096 instead of 2048. Customers may now deploy virtual machines with up to 256 disks using PVSCSI adapters. Each PVSCSI adapter can support up to 64 devices. Devices can be virtual disks or RDMs. A major change in 6.7 is the increased number of LUNs supported for Microsoft WSFC clusters. The number increased from 15 disks to 64 disks per adapter, PVSCSI only. This changes the number of LUNs available for a VM running MICROSOFT WSFC from 45 to 192 LUNs.

The increased maximums for Virtual SCSI adapter (PVSCSI only). Up to 64 Virtual SCSI Targets Per Virtual SCSI Adapter and up to 256 Virtual SCSI Targets Per Virtual Machine.

VMFS-3 EOL. Starting with vSphere 6.7, VMFS-3 will no longer be supported. Any volume/datastore still using VMFS-3 will automatically be upgraded to VMFS-5 during the installation or upgrade to vSphere 6.7. Any new volume/datastore created going forward will use VMFS-6 as the default.

Support for PMEM /NVDIMMs. Persistent Memory or PMem is a type of non-volatile DRAM (NVDIMM) that has the speed of DRAM but retains contents through power cycles. It’s a new layer that sits between NAND flash and DRAM providing faster performance and it’s non-volatile unlink DRAM.

Intel VMD (Volume Management Device). With vSphere 6.7, there is now native support for Intel VMD technology to enable the management of NMVe drives. This technology was introduced as an installable option in vSphere 6.5. Intel VMD currently enables hot-swap management, as well as NVMe drive, LED control allowing similar control used for SAS and SATA drives.

RDMA (Remote Direct Memory Access) over Converged Ethernet (RoCE). This release introduces RDMA using RoCE v2 support for ESXi hosts. RDMA provides low latency, and higher-throughput interconnects with CPU offloads between the end-points. If a host has RoCE capable network adaptor(s), this feature is automatically enabled.

Para-virtualized RDMA (PV-RDMA). In this release, ESXi introduces the PV-RDMA for Linux guest OS with RoCE v2 support. PV-RDMA enables customers to run RDMA capable applications in the virtualized environments. PV-RDMA enabled VMs can also be live migrated.

iSER (iSCSI Extension for RDMA). Customers may now deploy ESXi with external storage systems supporting iSER targets. iSER takes advantage of faster interconnects and CPU offload using RDMA over Converged Ethernet (RoCE). We are providing iSER initiator function, which allows ESXi storage stack to connect with iSER capable target storage systems.

SW-FCoE (Software Fiber Channel over Ethernet). In this release, ESXi introduces software-based FCoE (SW-FCoE) initiator than can create FCoE connection over Ethernet controllers. The VMware FCoE initiator works on lossless Ethernet fabric using Priority-based Flow Control (PFC). It can work in Fabric and VN2VN modes. Please check VMware Compatibility Guide (VCG) for supported NICs.

Performance

vSphere 6.7 VCSA delivers phenomenal performance improvements (all metrics compared at cluster scale limits, versus vSphere 6.5):
  • 2X faster performance in vCenter operations per second
  • 3X reduction in memory usage
  • 3X faster DRS-related operations (e.g. power-on virtual machine)

Security

vSphere 6.7 adds support for Trusted Platform Module (TPM) 2.0 hardware devices and also introduces Virtual TPM 2.0, significantly enhancing protection and assuring integrity for both the hypervisor and the guest operating system.

vSphere 6.7 introduces support for the entire range of Microsoft’s Virtualization Based Security technologies aka “Credential Guard” support.

Recoverability

vCenter Server Appliance (VCSA) File-Based Backup introduced in vSphere 6.5 now has a scheduler. Now customers can schedule the backups of their vCenter Server Appliances and select how many backups to retain. Another new section for File-Based backup is Activities. Once the backup job is complete it will be logged in the activity section with detailed information. The Restore workflow now includes a backup archive browser. The browser displays all your backups without having to know the entire backup path.

Conclusion

It seems that vSphere 6.7 is the continuous evolution of the best x86 virtualization platform with a lot of interesting improvements, features, and capabilities. Keep in mind, that this is just a list of features and capabilities which have to be very carefully planned, designed and tested before implementation into production.

Just FYI, I did not finish the reading of all vSphere 6.7 documents so I will update this blog post when find something interesting.

Friday, March 23, 2018

Deploying vCenter High Availability with network addresses in separate subnets

VMware vCenter High Availability is a very interesting feature included in vSphere 6.5. Generally, it provides higher availability of vCenter service by having three vCenter nodes (active/passive/witness) all serving the single vCenter service.

This is written in the official vCenter HA documentation
vCenter High Availability (vCenter HA) protects vCenter Server Appliance against host and hardware failures. The active-passive architecture of the solution can also help you reduce downtime significantly when you patch vCenter Server Appliance.
After some network configuration, you create a three-node cluster that contains Active, Passive, and Witness nodes. Different configuration paths are available.   
The last sentence is very true. The simplest VCHA deployment is within the same SSO domain and within the single datacenter with two Layer2 networks, one for management and second for the heartbeat. Such design can be deployed in fully automated manner and you just need to provide dedicated network (portgroup/VLAN) for the heartbeat network and use 3 IP addresses from separated heartbeat subnet. Easy. But is it what you are expecting from vCenter HA? To be honest, the much more attractive use case is to spread vCenter HA nodes across three datacenters to keep vSphere management up and running even one of two datacenters experiences some issue. Conceptually it is depicted in the figure below.

Conceptual vCenter HA Design
In this particular concept, I have embedded PSC controllers because of simplicity and vCenter HA can increase availability even of PSC services. The most interesting challenge in this concept is networking so let's look into the intended network logical design.

vCenter HA - networking logical design
Networking logical design:

  • Each vCenter Server Appliance node has two NICs
  • One NIC is connected to management network and second NIC to heartbeat network
  • Layer 2 Management network (VLAN 4) is stretched across datacenters A and B because vCenter IP address must work without human intervention in datacenter B after VCHA fail-over.
  • In each datacenter we have independent heartbeat network (VCHA-HB-A, VCHA-HB-B, VCHA-HB-C) with different IP subnets to not stretch Layer 2 across datacenters, especially not to datacenter C where is the witness. This requires specific static routes in each vCenter Server Appliance node to have IP reachability over heartbeat network.
  • Specific VCHA network tcp/udp ports must be allowed among VCHA nodes across a heartbeat network.
Helpful documents:

Implementation Notes: 

Note 1:
VMware KB 2148442 (Deploying vCenter High Availability with network addresses in separate subnets) is very important to deploy such design but one information is missing there. After cloning of vCenter Server Appliances, you have to go to passive node and configure on eth0 the same IP address you use in active node. Configuration is in file  /etc/systemd/network/10-eth0.network.manual
    Note 2:
    In case of badly destroyed VCHA cluster use following commands to destroy VCHA from the command line
    cd /etc/systemd/network 
    mv 10-eth0.network.manual 20-eth0.networkdestroy-vchareboot
      The solution was found at https://communities.vmware.com/thread/552084
        Link to the official documentation (Resolving Failover Failures) - https://docs.vmware.com/en/VMware-vSphere/6.5/com.vmware.vsphere.avail.doc/GUID-FE5106A8-5FE7-4C38-91AA-D7140944002D.html

        Note 3:
        In case, you will see the error message MethodFault.summary error during the finalization process it is because a hostname mismatch is detected. The hostname assigned to the Passive node must be the same as the hostname of the Active node. The solution was found  at https://www.altaro.com/vmware/how-to-deploy-a-vcenter-ha-cluster-part-2/ but also written in KB https://kb.vmware.com/kb/2148442

        Friday, March 09, 2018

        How to check I/O device on VMware HCL

        VMware has Hardware Compatibility List of supported I/O devices is available here
        https://www.vmware.com/resources/compatibility/search.php?deviceCategory=io

        VMware HCL for I/O devices

        The best identification of I/O device is VID (Vendor ID), DID (Device ID), SVID (Sub-Vendor ID), SSID (Sub-Device ID). VID, DID, SVID and SSID can be simply entered into VMware HCL and you will find if it is supported and what capabilities have been tested. You can also find supported firmware and driver.

        The get these identifiers you have to log in to ESXi via SSH and use command "vmkchdev -l".  This command shows VID:DID SVID:SSID for PCI devices and you can use grep to filter just VMware NICs (aka vmnic)
        vmkchdev -l | grep vmnic
        You should get similar output


        [dpasek@esx01:~] vmkchdev -l | grep vmnic
        0000:02:00.0 14e4:1657 103c:22be vmkernel vmnic0
        0000:02:00.1 14e4:1657 103c:22be vmkernel vmnic1
        0000:02:00.2 14e4:1657 103c:22be vmkernel vmnic2
        0000:02:00.3 14e4:1657 103c:22be vmkernel vmnic3
        0000:05:00.0 14e4:168e 103c:339d vmkernel vmnic4
        0000:05:00.1 14e4:168e 103c:339d vmkernel vmnic5
        0000:88:00.0 14e4:168e 103c:339d vmkernel vmnic6
        0000:88:00.1 14e4:168e 103c:339d vmkernel vmnic7

        So, in case of vmknic4 there is
        ·       VID:DID SVID:SSID
        ·       14e4:168e 103c:339d


        The same applies to HBAs and disk controllers.  For HBA and local disk controllers use
        vmkchdev -l | grep vmhba
        This is the output from my Intel NUC at my home lab

        [root@esx02:~] vmkchdev -l | more
        0000:00:00.0 8086:0a04 8086:2054 vmkernel 
        0000:00:02.0 8086:0a26 8086:2054 vmkernel 
        0000:00:03.0 8086:0a0c 8086:2054 vmkernel 
        0000:00:14.0 8086:9c31 8086:2054 vmkernel vmhba32
        0000:00:16.0 8086:9c3a 8086:2054 vmkernel 
        0000:00:19.0 8086:1559 8086:2054 vmkernel vmnic0
        0000:00:1b.0 8086:9c20 8086:2054 vmkernel 
        0000:00:1d.0 8086:9c26 8086:2054 vmkernel 
        0000:00:1f.0 8086:9c43 8086:2054 vmkernel 
        0000:00:1f.2 8086:9c03 8086:2054 vmkernel vmhba0

        0000:00:1f.3 8086:9c22 8086:2054 vmkernel 

        vmhba0 is local disk controller
        vmhba32 is USB storage controller
        vmnic0 is network interface

        Hope this helps.

        Wednesday, January 24, 2018

        vSphere 6.5 - DRS CPU Over-Commitment

        A lot of VMware vSphere architects and engineers are designing their vSphere clusters for some overbooking ratios to define some level of the service (SLA or OLA) and differentiate between different compute tiers. They usually want to achieve something like
        • Tier 1 cluster (mission-critical applications) - 1:1 vCPU / pCPU ratio
        • Tier 2 cluster (business-critical applications) - 3:1 vCPU / pCPU ratio
        • Tier 3 cluster (supporting applications) - 5:1 vCPU / pCPU ratio
        • Tier 4 cluster (virtual desktops) - 10:1 vCPU / pCPU ratio
        Terminology: 
        • vCPU - virtual CPU configured for Virtual Machine
        • pCPU - physical CPU core (hyper-threading is not considered) 
         
        Before vSphere 6.5 we have to monitor it externally by vROps or some other monitoring tool. Some time ago I have blogged how to achieve it with PowerCLI and LogInsight - ESXi host vCPU/pCPU reporting via PowerCLI to LogInsight.

        vSphere 6.5 DRS has introduced additional option to set maximum CPU over-commitment. It limits the number of vCPUs per pCPU in particular DRS cluster. However, it is good to know that there are two different advanced DRS options (configuration parameters) how to specify vCPU:pCPU ratio and each setting behaves differently. See table below ...

        DRS Advanced OptionScopeMin-Max Value
        MaxVcpusPerClusterPctcluster0% - 500%
        MaxVCPUsPerCorehost0 - 32

        It is worth to mention a little bit tricky setting of these additional options via GUI. It is good to know how GUI setting of "CPU Over-Commitment" is mapped to DRS cluster advanced options.

        In my lab, I have VCSA 6.5 U1c (build 7119157) so I did some tests.

        If I set "CPU Over-Commitment" in vSphere Web Client (Flash/Flex) it sets MaxVcpusPerClusterPct so it is the setting per the whole vSphere Cluster.

        However, in vSphere Client (HTML5) it sets MaxVCPUsPerCore so it is per ESXi host.

        Therefore, it is good to know what you would like to achieve and double check DRS Advanced Options.

        See different behavior in screenshots below

        6.5 U1c (build 7119157) vSphere Web Client (Flash/Flex) sets MaxVcpusPerClusterPct

        6.5 U1c (build 7119157) vSphere Client (HTML5) sets MaxVCPUsPerCore
        So this is how it should work. Now let's do some test to understand real behavior.

        TEST 1: MaxVcpusPerClusterPct = 0 

        Let's set MaxVcpusPerClusterPct to 0 so we are saying to allow 0 : 1 vCPU / pCPU ratio. In other words, no VM can be running in the cluster. And it works as expected. When I try to run VM I get the error "The total number of virtual CPUs present or requested in virtual machines' configuration has exceeded the limit on the host: 0". Well, it is a little bit misleading because it should be cluster-wide rule but it works as expected.

        Error when MaxVcpusPerClusterPct is set to 0
        TEST 2: MaxVcpusPerClusterPct = 100 

        Let's set MaxVcpusPerClusterPct to 100% so we are saying to allow 1 : 1 vCPU / pCPU ratio. I have 4 node DRS cluster where each ESXi host has two cores (pCPUs), therefore I have 8 pCPUs available in the cluster. And I can really start only four VMs because each has 2 vCPUs so I can run up to 8 vCPUs in DRS cluster.

        Error when MaxVcpusPerClusterPct is set to 100
        It is great, but it is worth to mention that VMs were started on single ESXi host even I have 4 ESXi hosts in DRS cluster. So vCPU / pCPU ratio is compliant per cluster but not per ESXi host as I have 8 vCPUs on single ESXi hosts having just 2 pCPUs. But that's expected behavior. So far so good.

        TEST 3: MaxVcpusPerCore = 0 

        MaxVcpusPerCore should solve the problem observed in previous two test because it should set vCPU / pCPU ratio per host which is much better from the predictability point of view.

        Let's set MaxVcpusPerCore to 0. My expectation was that I will not be able to start any VM but that was NOT the case. I was able to start a lot of VMs and exceed the expected vCPU / pCPU ratio. This is unexpected behavior.

        TEST 4: MaxVcpusPerCore = 1 


        Let's set MaxVcpusPerCore to 1.  My expectation was that I will not be able to start more than 2 vCPUs per ESXi host so only one VM with 2 vCPUs. Unfortunately, I was able to start much more vCPUs per ESXi host. This is again unexpected behavior.

        TEST 5: MaxVcpusPerCore = 4 

        I have been informed by DRS Engineer that the minimum value allowed by host over-commitment ratio option (MaxVCPUsPerCore) is 4:1.

        So, let's set MaxVcpusPerCore to 4.  I have prepared nine VMs with 4 vCPUs each and my expectation is that I will not be able to start more than 8 vCPUs per ESXi host so only two VMs with 4 vCPUs per ESXi host. And because I have 4 ESXi hosts per cluster I should be able to start the maximum of 8 VMs.

        Expected error message when MaxVcpusPerCore is set to 4 and vCPU:pCPU is over 4:1 per ESXi host.
        UPDATE 2018-04-20: The issue with MaxVcpusPerCore is fixed in vSphere 6.5 U2. This is written in Release Notes: The advanced vSphere DRS parameter MaxVcpusPerCore might not work as expected and the desired ratio of virtual CPUs per physical CPU or core will not take effect in configurations below 4:1. MaxVcpusPerCore supported ratios now start from 1:1.
        Please note, that MaxVcpusPerCore does not support ratio 0:1 in contrast with MaxVcpusPerClusterPct where 0:1 is possible and it will effectively disable to run any VM on the cluster.

        Conclusion

        Advanced DRS setting MaxVcpusPerClusterPct supports values between 0 and 500 and represents percentage between vCPUs and pCPUs across the whole DRS cluster. If the value is higher then 500, enforcing does not work. So, vCPU / pCPU percentage ratio can be enforced between 0:1 to 5:1.  vCPU:pCPU ratio 0:1 is a special setting where no VM can be PowerOn on vSphere Cluster. This is little bit risky setting but it can be used to put the whole cluster in kind of "maintenance mode" and forbid anyone to run VMs there.

        Advanced DRS setting MaxVcpusPerCore currently supports values between 0 and 32 but it works only with values between 4 and 32. This setting enforces vCPU / pCPU per each ESXi host within DRS cluster. This means that the minimum vCPU / pCPU ration configurable by this option is 4 : 1.

        To be honest, I think for vSphere architects/designers MaxVcpusPerCore makes more sense than MaxVcpusPerClusterPct because vCPU/pCPU overbooking ratio defines CPU quality per ESXi host.

        VMware is internally considering to align MaxVcpusPerClusterPct and MaxVcpusPerCore behavior and allow lower vCPU / pCPU ratios when MaxVcpusPerCore is used. I will track how this topic will evolve in the future. Stay tuned.

        Tuesday, September 19, 2017

        What is the difference between VMware vRealize Suite and vCloud Suite

        Several times I have been asked by my customers what is the difference between VMware vRealize Suite and vCloud Suite. Both are actually licensing packaging suits. VMware vCloud Suite suite is the superset of VMware vRealize Suite. In other words, vCloud Suite includes everything as vRealize Suite plus vSphere Infrastructure (ESXi Enterprise Plus licenses).

        VMware vRealize Suite is a purpose-built management solution for the heterogeneous data center and the hybrid cloud. It is designed to deliver and manage infrastructure and applications to increase business agility while maintaining IT control. It provides the most comprehensive management stack for private and public clouds, multiple hypervisors, and physical infrastructure.

        vRealize Suite editions comparison is available here and visualized on figure below.

        vRealize Suite editions comparison
        More specifically vRealize Suite 2017 includes following components:
        • vSphere Replication 6.5.1
        • vSphere Data Protection 6.1.5
        • vSphere Big Data Extensions 2.3.2
        • vRealize Orchestrator Appliance 7.3.0
        • vRealize Suite Lifecycle Manager 1.0
        • vRealize Operations Manager 6.6.1
        • vRealize Business for Cloud 7.3.1
        • vRealize Log Insight 4.5.0
        • vRealize Automation 7.3.0
        • vRealize Code Stream Management Pack for IT DevOps 2.2.1
        VMware vCloud Suite brings together VMware’s industry-leading vSphere hypervisor with vRealize Suite, the complete cloud management solution.

        vCloud Suite 2017 currently includes vRealize Suite 2017 and vSphere Enterprise Plus and NSX-V:
        • Everything included in vRealize Suite 2017
        • plus vSphere Enterprise Plus Infrastructure for vCloud Suite (ESXi licenses only, vCenter is not included)
        • ESXi 6.5 U1
        • plus NSX-V
        • NSX-V 6.3.3
        What's new in 2017 Suites?
        1. VMware has introduced new overall manager called vRealize Suite Lifecycle Manager to simplify deployment and on-going management of the vRealize products.
        2. NSX-V (NSX for vSphere) is included in vCloud Suite 2017 ???
        Please note, that vCenter Server is not included in vCloud Suite, therefore additional license for vCloud Suite is required.

        Hope this helps to understand VMware license packaging.

        ...........................................................................................

        UPDATE 2017-09-20:  The note about NSX-V licensing in vCloud Suite 2017 Release announcement is a little bit misleading. On VMware WebSite here is written
        NSX is available as an optional component that can be purchased with vCloud Suite. 
        So, VMware vCloud Suite 2017 customers are not entitled to use full NSX-V. It is just an entitlement to use NSX Manager for AV (replacement for vCNS Endpoint). NSX Manager for AV has been available for free even before vCloud Suite 2017 so that's nothing new. And of course, any VMware customer can purchase full NSX-V to extend their vSphere infrastructure for a network virtualization and datacenter modernization.

        Tuesday, August 15, 2017

        NSX Basic Concepts, Tips and Tricks

        NSX and Network Teaming

        There are multiple options how to achieve network teaming from ESXi to the physical network. For more information see my another blog post "Back to the basics - VMware vSphere networking".

        In a nutshell, there are generally three supported methods how to connect NSX VTEP(s) to the physical network
        1. Explicit failover - only single physical NIC is active at any given time, therefore no load balancing at all
        2. LACP - single aggregated virtual interface where load balancing is done based on hashing algorithm
        3. Switch independent teaming achieved by multiple VTEPs where each VTEP is bind to different ESXi pNIC.
        Let's assume we have switch independent teaming with multiple independent uplinks to the physical network. Now the question is how to check VM vNIC to ESXi host pNIC mapping? I'm aware of at least four methods how to check this mapping
        1. ESXTOP
        2. ESXCLI
        3. NSX Controller
        4. NSX Manager
        1/ ESXTOP method
        • ssh to ESXi
        • run esxtop
        • Press key [n] to switch to network view
        • Check column TEAM-PNIC – it should be different vmnic (ESXi pNIC) for each VM
        2/ ESXCLI method
        • ssh to ESXi
        • Use command “esxcli network vm list” and locate World IDs of VM
        • Use “esxcli network vm port list -w ” and check “Team Uplink” value. It should be different vmnic (ESXi PNIC) for each VM
        3/ NSX Controller method
        • Identify MAC address of VM
        • Login to NSX Controller nodes (ssh or console) one by one
        • Use command “show control-cluster logical-switches mac-table ” to show mac-address to VTEP mappings. I assume multi VTEP configuration where each VTEP is statically bound to particular ESXi pNIC (vmnic)
        4/ NSX Manager method
        • Identify MAC address of VM
        • Login to NSX Manager (ssh or console)
        • Go through all controllers and show mac address table where is also information behind which VTEP particular mac address is
        • i) show controller list all
        • ii) show logical-switch controller controller-1 vni 10001 mac
        • iii) show logical-switch controller controller-2 vni 10001 mac
        • iv) show logical-switch controller controller-3 vni 10001 mac
        The appropriate method is typically chosen based on the role and Role Based Access Control. vSphere Administrator will probably use esxtop or esxcli and Network Administrator will use NSX Manager or Controller.

        Distributed Logical Router (DLR)

        DLR is a virtual router distributed across multiple ESXi hosts. You can imagine it as a chassis with multiple line cards.  Chassis is virtual (software based) and line cards are software modules spread across multiple ESXi hosts (physical x86 servers).

        The basic concept of DLR is that every routing decision is done locally which means that NSX DLR always performs local routing on the DLR instance running in the kernel of the ESXi hosting the workload that initiates the communication. When VM traffic needs to be routed to another logical switch, it first comes to DLR on the same ESXi host where VM is running. Each DLR line card module (ESXi host) has all logical switches (VXLANs) connected locally so DLR forwards the packet to the appropriate destination logical switch and if the target VM runs on another ESXi host the packet is encapsulated on local ESXi host and decapsulated on target ESXi host.

        It is good to know, that DLR uses always the same MAC address for default gateway addresses for all logical switches. This MAC address is called VMAC. This is a MAC address used for DLR logical L3 interfaces (LIFs) connected into logical switches (VXLANs).

        However, there must be some coordination between multiple DLR "line card" modules (ESXi hosts) therefore each DLR module must also have physical MAC address. This MAC address is called PMAC.

        To show DLR PMAC and VMAC run following command on ESXi host
        net-vdr -l -C

        Distributed Logical Firewall (DFW) - firewall rules

        NSX Distributed Firewall applies firewall rules directly to VM vNICs. In the vNIC is the concept of slots where different services are bind and chain together. NSX DFW sits in slot 2 and for example, the third party firewall sits in slot 4.

        So the DFW firewall rules are automatically applied on each vNIC so the question is how to double check what rules are at vNIC level.

        There are two methods how to check it
        1. ESXi commands
        2. NSX Manager commands
        1/ ESXi method
        • ssh to ESXi
        • Use command “summarize-dvfilter” and locate the VM of your interest and its vNIC name is slot 2 used by agent vmware-sfw
        • grep commands can help us here ... "summarize-dvfilter | grep -A 10 "
        •  vNIC name should looks similar to nic-24565940-eth0-vmware-sfw.2
        • Now you can list firewall rules by command "vsipioctl getfwrules -f nic-24565940-eth0-vmware-sfw.2"

        2/ NSX Manager method (https://kb.vmware.com/kb/2125482)
        • Log in to the NSX Manager with the admin credentials
        • To display a summary of DVFilter information, run the command "show dfw host-id summarize-dvfilter"
        • To display detailed information about a vnic, run the command "show dfw host host-id vnic"
        • To display the rules configured on the filter, run the command "show dfw host host-id vnic vnic-id filter filter-name rules"
        • To display the addrsets configured on the filter, run the command "show dfw host host-id vnic vnic-id filter filter-name addrsets"
        And again, the appropriate method is typically chosen based on the administrator role and Role Based Access Control. 

        Distributed Logical Firewall (DFW) - third party integration and availability considerations

        NSX Distributed Firewall supports integration with third party solutions. This integration is also called service chaining. Third party solution is hooked to a particular vNIC slot and usually, some selected or potentially all (not recommended) traffic can be redirected to third-party solution agent running on each ESXi host as a special Virtual Machine. The third-party solution can inspect the traffic and allow or deny the traffic. However,  what happens when agent VM is not available? It is easy to test it, you can Power Off Agent VM and see what happens. Actually, the behavior depends on Service failOpen/failClosed policy.  You can check policy setting as depicted on the screenshot below ...

        Service failOpen/failClosed policy
        If failOpen is set to false then the virtual machine traffic will be dropped in case the agent is unavailable. It has a negative impact on availability but positive impact on security. If failOpen is set to true then the VM traffic will be allowed and everything works even the agent is not available. In such situation, the security policy cannot be enforced and there is a potential security risk. So this is typical design decision point where a decision is dependent on customer specific requirements.

        Now the question is how failOpen setting can be changed. Well, my understanding is that it depends on third party solution. Here is the link to TrendMicro how to - "Set vNetwork behavior when appliances shut down"  

        Sunday, June 25, 2017

        Start order of software services in VMware vCenter Server Appliance 6.0 U2

        vCenter Server Appliance 6.0 U2 services are started in the following order ...

        1. vmafdd (VMware Authentication Framework)
        2. vmware-rhttpproxy (VMware HTTP Reverse Proxy)
        3. vmdird (VMware Directory Service)
        4. vmcad (VMware Certificate Service)
        5. vmware-sts-idmd (VMware Identity Management Service)
        6. vmware-stsd (VMware Security Token Service)
        7. vmware-cm (VMware Component Manager)
        8. vmware-cis-license (VMware License Service)
        9. vmware-psc-client (VMware Platform Services Controller Client)
        10. vmware-sca (VMware Service Control Agent)
        11. applmgmt (VMware Appliance Management Service)
        12. vmware-netdumper (VMware vSphere ESXi Dump Collector)
        13. vmware-syslog (VMware Common Logging Service)
        14. vmware-syslog-health (VMware Syslog Health Service)
        15. vmware-vapi-endpoint (VMware vAPI Endpoint)
        16. vmware-vpostgres (VMware Postgres)
        17. vmware-invsvc (VMware Inventory Service)
        18. vmware-mbcs (VMware Message Bus Configuration Service)
        19. vmware-vpxd (VMware vCenter Server)
        20. vmware-eam (VMware ESX Agent Manager)
        21. vmware-rbd-watchdog (VMware vSphere Auto Deploy Waiter)
        22. vmware-sps (VMware vSphere Profile-Driven Storage Service)
        23. vmware-vdcs (VMware Content Library Service)
        24. vmware-vpx-workflow (VMware vCenter Workflow Manager)
        25. vmware-vsan-health (VMware VSAN Health Service)
        26. vmware-vsm (VMware vService Manager)
        27. vsphere-client ()
        28. vmware-perfcharts (VMware Performance Charts)
        29. vmware-vws (VMware System and Hardware Health Manager) 


        Thursday, June 22, 2017

        CLI for VMware Virtual Distributed Switch

        A few weeks ago I have been asked by one of my customers if VMware Virtual Distributed Switch (aka VDS) supports Cisco like command line interface. The key idea behind was to integrate vSphere switch with open-source tool Network Tracking Database (NetDB) which they use for tracking MAC addresses within their network. I have been told by customer that NetDB can telnet/ssh to Cisco switches and do screen scraping so would not it be cool to have the most popular switch CLI commands for VDS? These commands are

        • show mac-address-table
        • show interface status
        The official answer is NO, but wait a minute. Almost anything is possible with VMware API. So my solution is leveraging VMware's vSphere Perl SDK to pull information out of Distributed Virtual Switches. I have prepared PERL script vdscli.pl which currently supports two commands mentioned above. It goes through all VMware Distributed Switches on single vCenter.

        Script along with shell wrappers are available on GITHUB here https://github.com/davidpasek/vdscli
        See screenshots below to get an idea what script does.

        The output of the command
        vdscli.pl --server=vc01.home.uw.cz --username readonly --password readonly --cmd show-port-status
        looks as depicted in screenshot below.


        and output of the command
        vdscli.pl --server=vc01.home.uw.cz --username readonly --password readonly --cmd show-mac-address-table

        So now we have Perl scripts to get information from VMware Distributed Virtual Switch which is nice, however, we would like to have Interactive CLI to have the same user experience as we have on physical switches CLI, right? For Interactive CLI I have decided to use Python ishell (https://github.com/italorossi/ishell) to emulate Cisco like CLI. To start interactive VDSCLI shell you must have Python with iShell installed and then you can simply run script

        ./vdscli-ishell.py

        which is just a wrapper around vdscli.pl The screenshot of VDSCLI shell is in the figure below

        VDSCLI Interactive Shell
        And the last step is to allow SSH or Telnet access to VDSCLI shell. It can be very easily done via standard Linux possibility to change a shell for the particular user. The VDSCLI over ssh is depicted on the screenshot below.

        VDSCLI Interactive Shell over SSH
        To operationalize all these scripts, I would highly encourage you to read my another blog post ...
        "CLI for VMware Virtual Distributed Switch - implementation procedure".

        Hope somebody else in VMware community will find it useful.

        Tuesday, October 18, 2016

        vSphere 6.5 announced so what is coming?

        vSphere 6.5 has been announced on VMworld 2016 so you can ask yourself what it brings and why consider upgrade or at least upgrade plan.

        It is obvious and expected that almost all vSphere 6.5 scalability limits will be increased. Configuration maximums like hosts per vCenter, powered on VMs per vCenter, hosts per cluster, VMs per cluster, vCenters in linked mode, etc are expected to increase. Theses limits are no longer limits for me but if you need it, just wait for vSphere 6.5 GA and double check well known document vsphere-65-configuration-maximums.pdf

        However, vSphere users are usually looking for new features. So here they are ...

        vCenter features
        • vCenter Server Appliance (aka VCSA) will be recommended as "First Choice" because the coolest new features are available just in VCSA. 
        • Platform Service Controller (PSC) will have out-of-the-box high availability for VCSA using PSC's. You will be able to achieve PSC RTO lower then 5 minutes. << UPDATE: unfortunately, this feature was not released in vSphere 6.5 release so let's hope it will be released in future vSphere 6.5 Updates.  
        • VCSA supports native High Availability support of vCenter service with RTO lower then 15 minutes.
        • VCSA has embedded vCenter Appliance Monitoring and Management to gain visibility into VCSA performance and capacity management including embedded vPostgreSQL database service.
        • VMware Update Manager (VUM) is be fully integrated into VCSA.
        • File level vCenter Server Backup and Restore is complementary backup method to existing VDP image backup. It will be possible to restore vCenter file level backup to fresh VCSA.
        • Content Libraries in vSphere 6.5 have additional features including the option to mount an ISO from a Content Library, update existing templates, and apply guest OS Customization Specifications during VM deployments. If Content Libraries reside on VCSA then you can also make use of vCenter HA, and native Backup and Restore, both new features to vSphere 6.5 mentioned above.
        vSphere HA Cluster features
        • vSphere HA cluster wide restart ordering ability with intra-app dependencies during failover. It allows multi-tier application consistency during VM fail-overs.  It is also known as "vSphere HA Orchestrated Restart" because you can create VM to VM dependencies which will force specified VMs to perform HA restarts before others. You can also choose in your vSphere HA settings, when the next VM should begin restarting. At the power-on initiated command, when resources allocated, VMware Tools heartbeats,  etc. You can also set additional timeouts and delays if needed.
        • vSphere HA Admission Control - default Admission Control policy has changed from Slot Policy (Default until 6.5), to ‘Cluster Resource Percentage’. Any time you add or remove a host from the cluster, the failover capacity percentages will update, and the amount of resources required on each host will also be updated automatically.
        • Proactive HA - it integrates with the Server vendor’s monitoring software, via a Web Client plugin, which will pass detailed server health status/alerts to DRS, and DRS will react based on the health state of the host’s hardware. Yes, even the name is "Proactive HA" it is DRS functionality. Confusing? The name was chosen because it has positive impact on availability.
        vSphere DRS Cluster features
        • Predictive DRS - it integrates DRS with vROps to provide placement and balancing decisions.
        • Network-Aware DRS - DRS takes physical NIC utilization in to consideration. Once a target host has been chosen for placement/load-balancing, DRS will then check to see if that host’s network is saturated (default is 80% utilization of connected uplinks, but can be configured with ‘NetworkAwareDrsSaturationThresholdPercent’. If the host is considered saturated, it will use a different target host
        • DRS Additional Option : VM Distribution - even distribution of VMs across cluster
        • DRS Additional Option : Memory Metric for Load Balancing - usage of active versus consumed memory for DRS recommendations
        • DRS Additional Option : CPU Over-Commitment - limit the number of vCPUs per pCPU in particular DRS cluster. Specific vCPU:pCPU ratio is set as advanced DRS option MaxVcpusPerClusterPct.
        ESXi features and improvements
        • ESXi is pretty stable and best in class hypervisor. However even in this component you can expect some improvements.  For example I/O improvements because of RDMA / PVRDMA. PVRDMA (para-virtualized RDMA) is industry first virtualized RDMA and it allows virtualization of applications which require ultra low latency. And it supports live vMotion which SR-IOV does not.
        • ESXi core storage improvements - Support for 4K Native Drives in 512e mode, SE Sparse Default for VMFS, Automatic Space Reclamation, Support for 1024 devices and 4096 paths (versus 256 and 1024 in the previous versions)
        vSphere management features
        • Auto Deploy and Image Builder will be full integrated into WebClient and Host profiles will be improved to smoothly support auto deploy.
        • vSphere Web Client usability and performance will be improved again. It is pretty important because C# client is not available for vSphere 6.5 so vSphere admins will rely on web client. HTML5-based vSphere Client should be included in 6.5 release.
        • Content library improvements - mount ISO directly from content library, customization during VM deployment, improved scale and performance, high availability along with VCSA
        • vSphere 6.5 introduces new REST-based APIs for VM Management
        Storage related features
        • VVOLs 2.0 will bring data protection and replication along with support for MSCS, Oracle RAC, NFS 4.1 and SMP-FT.
        • VSAN - Virtual SAN iSCSI Service
        • VSAN - 2-Node Direct Connect with witness Traffic Separation for ROBO
        • VSAN - 512e drive support (still waiting for 4K native support)
        Security related features
        • VM Encryption will be new feature to protect your VM data with tenants keys. It enables encryption on a per VM as well as per VMDK basis. It can be integrated with 3rd party Key Management Servers (KMS).
        • vSphere 6.5 also delivers enhanced audit-quality logging capabilities that provide more forensic information about user actions.
        WebClient related features
        • In vSphere 6.5, the vSphere Web Client will have no dependency on Client Integration Plug-in (as it exists before).  For the Use Window Session Authentication functionality, you will need the new slimmed down Enhanced Authentication Plug-in, but the other functions (File upload/download, Deploy OVA/OVF) are replicated without CIP.
        Conclusion

        vSphere 6 is already very mature virtualization platform but vSphere 6.5 brings some very interesting enterprise features if you ask me. The most interesting features for me personally are
        • VCSA and PSC high availability
        • VVOLs 2.0
        • VM Encryption
        • REST-based APIs for vSphere Management
        but all other features are cool and very handy as well. 

        It is very common practice to wait for Update 1 before upgrading production environments but our labs and test environments are good candidates for vSphere 6.5 release when available. I'm eagerly waiting for GA.

        Other related blog posts and resources:

        Monday, October 17, 2016

        VMware SIOC quick configuration in datacenter scale

        I'm currently troubleshooting one weird high kernel latency (KAVG) issue and there is a suspicion that the issue can be somehow related to VMware SIOC which is widely use in customer's environment. To confirm or disprove the issue is really related to SIOC we can simply disable SIOC on all datastores and observe if it has positive impact on kernel latency.

        Customer has lot of production datastores grouped in datastore clusters so following PowerCLI one liners can help with quick configuration and validation of SIOC settings across whole datacenter.

        SIOC current state for all datastores in datastore clusters
         Get-DatastoreCluster | Get-Datastore | select-object name,type,StorageIOControlEnabled | Format-List -Property *  
        

        Disable SIOC for all datastores in datastore clusters
         Set-Datastore (Get-DatastoreCluster | Get-Datastore) -StorageIOControlEnabled $false | select-object name,type,StorageIOControlEnabled
        

        Enable SIOC for all datastores in datastore clusters
         Set-Datastore (Get-DatastoreCluster | Get-Datastore) -StorageIOControlEnabled $true | select-object name,type,StorageIOControlEnabled  
        

        Thanks PowerCLI!

        Thursday, October 13, 2016

        Metro Cluster High Availability or SRM Disaster Recovery?

        Several years I continuously try to explain my customers that metro cluster is not disaster recovery. I have finally found some time and summarize my thoughts into slide deck which I published on SlideShare. I'm planning to present it at Czech VMUG local meeting on 6 December this year. More info about this particular Czech VMUG event is here.

        The goal of my presentation is to explain the difference between multi site high availability (aka metro cluster) and disaster recovery. General concepts are same for any products but presentation is obviously more tailored for specific VMware products and technologies.

        You can look at presentation here on SlideShare ...



        It would be great to see you at the event if you will be in the town. But in the meantime don't hesitate to write any comment or feedback here and we can have good discussion as there are still two months till the event.

        BTW: Kudos to Stanislav Jurena @stan_jurena who already did several reviews and gave me some comments and feedback before first public release.  

        Thursday, May 19, 2016

        VMware vSphere SDRS - test plan of SDRS initial placement

        VMware vSphere Storage DRS (aka SDRS) stands for Storage Distributed Resource Scheduler. It continuously balances storage space usage and storage I/O load while avoiding resource bottlenecks to meet application service levels.

        Lab environment:
        5x10GB Datastores formed into Datastore Cluster with SDRS enabled.
        It is configured to balance based on storage space usage and also I/O load.

        • Storage Space threshold is 1 GB
        • I/O latency threshold is kept on default 15 ms.
        • Each "empty" datastore has real capacity 9.75 GB where real free capacity is 8.89 GB because 882 MB is used. 

        You can see configuration details on screenshot below.  


        Capacity of one particular 10 GB datastore is depicted below.
        Used space (882 MB) is occupied by following system files (.sf) ...


        In VMFS 5 every datastore gets its own hidden files to save the file-system structure.

        Test 1

        Test description: Does SDRS Initial Placement algorithm take into account VM swap file capacity?

        Test prerequisites:
        • All 5 datastores in datastore cluster are empty
        • That's mean that each datastore has free capacity 8.89 GB
        • Provisioned VM doesn't have any RAM reservation
        Test steps:
        • Deploy Virtual Machine with 4 GB RAM and 8 GB Disk manually (through Web Client)
        • Start deployed Virtual Machine
        • Observe behavior 
        Test expectations:

        • I want to test if swap file is considered during SDRS initial placement
        • We have only 8.89 GB free space on datatastores therefore if VM swap file is considered new VM with 8GB disk and 4 GB RAM wont be provisioned because we would need 12 GB space on some datastore which is not our case.
        • In other words, if provisioning fails then we will proof that SDRS doesn't take VM swap file into account.

        Test screenshots:

        Deployed Virtual Machine.
        VM PowerOn Failure 
        Test Result:

        • Virtual machine was successfully provisioned and 8GB was decreased from Datastore 5 available space.
        • Virtual machine power on action failed because of not enough storage space for 4 GB swap file. This is expected behavior in case that SDRS doesn't take VM swap into account.

        Test Summary:

        • We have tested that SDRS Initial Placement algorithm does NOT take VM swap file capacity into account.
        • Virtual Machine memory (RAM) reservation would have impact on such test because if VM has for example 100% memory reservation it doesn't need any disk space for VM swap.

        Test 2

        Test description: How SDRS defragmentation is efficient when datastore cluster is running out of storage space?

        Test prerequisites:
        • 4 datastores in datastore cluster are almost full
        • 1 datastore (Datastore4) has 7.58 GB free space
        • In one datastore (Datastore5) we have virtual machine (test1_big) having 8GB disk
        Test steps:
        • Clone Virtual Machine (test1_big) to datstore cluster (through Web Client)
        • Observe behavior 
        Test expectations:
        • SDRS will free up Datastore4 to have enough space for clone of virtual machine (test1_big)
        • Provisioning of virtual machine clone will be successful 

        Test screenshots:
        Before SDRS defragmentation
        After SDRS defragmentation and clone provisioning

        Test Result:
        • SDRS freed up Datastore4 as expected
        • Provisioning of virtual machine clone FAILED because of insufficient disk space on Datastore4. 
        • That's unexpected behavior because Datastore4 is empty (thanks to SDRS defragmentation) and another machine with same configuration was successfully provisioned on Datastore5.
        Test Summary:

        • SDRS successfully freed up the only datastore where virtual machine clone can be placed but VM clone deployment started before storage vMotion finished therefore clone provisioning failed.
        • SDRS defragmentation works but there can be some cases when initial placement fails even the storage was freed up and there will be free continuous space in some datastore after defragmentation.
        • It is important to understand how VM provisioning to datastore cluster really works. Datastore Cluster is nothing else then the group of single datastores where SDRS is "just" a scheduler on top of Datastore Cluster. You can imagine a scheduler as a placement engine which prepare placement recommendations for initial placement and continuous balancing. That means that other software component (C# Client, Web Client, PowerCLI, vRealize Automation, vCloud Director, etc) is responsible for initial placement provisioning and SDRS give them recommendations where is the best place to put a new storage objects (vmdk file or VM config file).
        • In other words, Initial VM provisioning doesn’t have nothing to do with SDRS initial placement. VM initial provisioning process is managed by vSphere Client, vRA, vRO, PowerCLI or other software component over vSphere API. SDRS is just a placement engine gives recommendation where is the best place at the moment when is asked for recommendations. Provisioning process selects one particular SDRS recommendation and continue with provisioning (API method ApplyStorageDrsRecommendation_Task). However, in the mean time there can be some other software doing VM provisioning and selected datastore can be filled by somebody else. There is always some probability for vm provisioning failure and it is exactly where good vSphere / Storage design has crucial role to decrease probability of provisioning failure. 

        Test 3

        Test description: How is SDRS initial placement balancing among different datastores?

        Test prerequisites:
        • Storage Space threshold is 1 GB
        • I/O latency threshold is kept on default 15 ms.
        • Each "empty" datastore has real capacity 9.75 GB where real free capacity is 8.89 GB because 882 MB is used. 
        • Usage of PowerCLI script to provision multiple VMs. PowerCLI script is available here.
        Test steps:
        • Run PowerCLI script to generate 50 virtual machines with following specification (1 vCPU, 512 MB RAM, 1GB Disk - thick) in to datastore cluster with SDRS enabled.
        • Observe behavior
        Test expectations:
        • We have datastore cluster with 5 datastores each having 8.89 GB (9,103 MB) available storage.
        • We are deploying VMs with 1000 MB each.
        • It is deployed in not power on state - so swap file doesn't need to be considered. 
        • Therefore we would expect to end up with 45 VMs balanced in round robin fashion across 5 datastores.  
        Test screenshots:
        Single datastore capacity
        PowerCLI Automated Provisioning.
        Datastore free space after automatic sequential provisioning

        Test Result:
        • 45 VMs was successfully provisioned and 46th-50th VM failed because of "Insufficient disk space on datastore 'Datastore1'." This was expected behavior.
        • Following VMs are provisioned on datastores
        • Datastore 1: TEST-05, TEST-06, TEST-15, TEST-20, TEST-21, TEST-30, TEST-31, TEST-40, TEST-45 
        • Datastore 2: TEST-04, TEST-10, TEST-11, TEST-19, TEST-25, TEST-26, TEST-35, TEST-36, TEST-44
        • Datastore 3: TEST-03, TEST-09, TEST-14, TEST-18, TEST-24, TEST-29, TEST-34, TEST-39, TEST-43
        • Datastore 4: TEST-02, TEST-08, TEST-13, TEST-17, TEST-23, TEST-28, TEST-33, TEST-38, TEST-42
        • Datastore 5: TEST-01, TEST-07, TEST-12, TEST-16, TEST-22, TEST-27, TEST-32, TEST-37, TEST-41
        Test Summary: Test passed as expected. Only few details are worth to mention.
        • I would expect VMs evenly distributed across datastores.  Recall that we are using artificial sequence provisioning of 1GB vDisks per VM. I would expect VMs TEST-01, TEST-06, TEST-11, TEST-16, TEST-21, TEST-26, TEST-31, TEST-36, TEST-41 on Datastore 5. And similar VM numbering on other datastores. But at the end of the day it doesn't seems to be a big deal.
        • Please, note that I observed that different provisioning runs can end-up with slightly different machine placement. I have suspicious that it is because other factors (I/O load, storage usage trend) then are also considered in SDRS algorithm.
        • Datastore free space is 103 MB on all datastores. Recall that we have Storage Space threshold set to 1 GB. That's expected behavior. Storage Space threshold is just a threshold (soft limit) used by SDRS for balancing and defragment. 
        Test 4

        Test description: Will be new VM provisioned to the datastore with the biggest frees space?

        Test prerequisites:
        • Storage Space threshold is 1 GB
        • I/O latency threshold is kept on default 15 ms.
        • One datastore (Datastore1) is "empty" has real capacity 9.5 GB where real free capacity is 8.64 GB because 882 MB is used. 
        • One datastore (Datastore2) has 5.71 GB free capacity.
        • All other datastores (Datastore3, Datastore4, Datastore5) are almost full having only 848 MB empty.
        Test steps:
        • Usage of vSphere Web Client to provision one VM with 2GB disk into Datastore Cluster.
        • Observe behavior. We are interested where new VM will be placed.
        Test expectations:
        • We expect that new VM disk will be placed on Datastore1 because there is the biggest free (available) space.

        Test screenshots:
        Datastore cluster capacity before VM provisioning.
        Datastore cluster capacity after VM provisioning.
        Test Result:
        • New virtual machine was provisioned into Datastore1 where was the bigest available storage capacity.
        Test Summary: Initial placement behaves as expected. New VM is placed to the datastore with less used space. However, we should be aware that this test was done just for single VM provisioning. Multiple VM provisioning can behaves differently because of other SDRS calculation factors (I/O load, capacity usage trend) and also because of particular provisioning workflow and exact timing when SDRS recommendation is called and when datastore space is really consumed for next SDRS recommendations.

        Next steps

        See blog post "Storage DRS Design Considerations".

        And as always, any comment is appreciated.