Disk requirements

Service-specific disk configurations are listed in Deployment scenarios. For known hardware limitations and compatibility considerations, refer to Hardware compatibility.

These requirements describe the disk roles, capacity, performance, and configuration requirements for Virtuozzo Infrastructure.

Requirements by disk role

The following requirements apply according to disk role.

Disk role Quantity Disk size Type Endurance
System One disk per node 120 GB minimum; 240-480 GB recommended, depending on node role SSD; RAID1 recommended for production
Metadata One metadata service per node; three to five metadata services per cluster 100 GB Enterprise-grade SSD with power loss protection. If possible, use a System+MDS layout instead of combining MDS with a cache role. 1 DWPD minimum
Cache Optional; one SSD per 4-12 HDDs 100+ GB Enterprise-grade SSD with power loss protection and 75 MB/s sequential write performance per serviced HDD 1 DWPD minimum; Mix-Use 3 DWPD or Write-Intensive 10 DWPD recommended
Storage At least one per storage node; minimum three per cluster 100 GB minimum; unlimited size on a physical server; 10 TB maximum recommended inside a virtual machine SATA/SAS HDD or enterprise-grade SATA/SAS/NVMe SSD with power loss protection 1 DWPD minimum

Disk types and roles

  • Using SATA HDDs with one SSD for caching is more cost effective than using only SAS HDDs without such an SSD.
  • Using NVMe or SAS SSDs for write caching improves random I/O performance and is highly recommended for all workloads with heavy random access (for example, iSCSI volumes). In turn, SATA disks are best suited for SSD-only configurations but not write caching.
  • Run metadata services on the system SSD by using the System+MDS role. On modern all-SSD or SSD-cached clusters, do not place metadata on a cache SSD, because the cache device is already under heavy load.
  • For the best performance, use NVMe SSDs with direct CPU attachment and avoid extra controllers or expanders in the data path.
  • Using shingled magnetic recording (SMR) HDDs is available only for storage purposes and only if the node has an SSD disk for cache.
  • If capacity is the main goal and you need to store infrequently accessed data, select SATA disks over SAS ones. If performance is the main goal, select NVMe or SAS disks over SATA ones.
  • Disk block size (for example, 512b or 4K) is not important and has no effect on performance.
  • The maximum supported physical partition size is 254 TiB.

Disk capacity

  • The system disk must have at least 120 GB of space. For production, use mirrored SSDs and size the system disk for service logs, swap, and the services hosted on the node.
  • It is possible to use disks of different size in the same cluster. However, keep in mind that, given the same IOPS, smaller disks will offer higher performance per terabyte of data compared to bigger disks. It is recommended to group disks with the same IOPS per terabyte in the same tier.
  • The capacity of HDD and SSD is measured and specified with decimal, not binary prefixes, so “TB” in disk specifications usually means “terabyte.” The operating system, however, displays a drive capacity using binary prefixes meaning that “TB” is “tebibyte” which is a noticeably larger number. As a result, disks may show a capacity smaller than the one marketed by the vendor. For example, a disk with 6 TB in specifications may be shown to have 5.45 TB of actual disk space in Virtuozzo Infrastructure. 5 percent of disk space is reserved for emergency needs. Therefore, if you add a 6 TB disk to a cluster, the available physical space should increase by about 5.2 TB.
  • Performance of SSD disks may depend on their size. Lower-capacity drives (100 to 400 GB) may perform much slower (sometimes up to ten times slower) than higher-capacity ones (1.9 to 3.8 TB). Check the drive performance and endurance specifications before purchasing hardware.
  • Thin provisioning is always enabled for all data and cannot be configured otherwise.

RAID and HBA controllers

  • Create hardware or software RAID1 volumes for system disks by using RAID or HBA controllers, respectively, to ensure its high performance and availability.
  • Use HBA controllers, as they are less expensive and easier to manage than RAID controllers.
  • Disable all RAID controller caches for SSD drives. Modern SSDs have good performance that can be reduced by a RAID controller’s write and read cache. It is recommended to disable caching for SSD drives and leave it enabled only for HDD drives.
  • If you use RAID controllers, do not create RAID volumes from HDDs intended for storage. Each storage HDD needs to be recognized by Virtuozzo Infrastructure as a separate device.
  • If you use RAID controllers with caching, equip them with backup battery units (BBUs), to protect against cache loss during power outages.

Disk layouts and operations

  • Each management node must have at least two disks: one for system and metadata, and one for storage. Each secondary node must have at least two disks: one for system and one for storage.
  • Use HDD-only, HDD plus system SSD, HDD plus SSD cache, or SSD-only layouts depending on the workload and cost target.
  • Cache devices can improve performance but must be sized and protected carefully. A failed cache device can make all capacity devices associated with it unavailable.
  • Use UPS devices for all servers and network switches, and use enterprise SSDs with power loss protection.
  • Before production, check that all devices used for CS journals, MDS journals, and chunk servers can flush cached data to disk during a power failure.

HDD/SSD configuration

Minimum disk count and role placement depend on the selected storage configuration, as shown below.

Configuration Best for Disk roles (per node)
HDD only Lowest cost 1 HDD system, 1 HDD metadata, remaining HDD storage
HDD + system SSD (no cache) Capacity-oriented 1 SSD system+metadata, remaining HDD storage
HDD + SSD cache Performance-oriented 1 SSD system+metadata, 1 SSD cache, remaining HDD storage
SSD only Highest performance (no cache needed) 1 SSD system+metadata, remaining SSD storage
Multi-tier Mixed workloads on one cluster Assign disk groups to tiers (e.g. NVMe hot tier + HDD-cached capacity tier)

Cache configuration

SSD cache devices improve the performance of HDD-based storage tiers and are not required for SSD-only storage.

  • Use one cache SSD for every 4–12 HDDs.
  • Provide at least 100 GB of cache capacity.
  • Use enterprise-grade SSDs with power loss protection.
  • Use SSDs with at least 1 DWPD endurance. For write-intensive workloads, use 3–10 DWPD SSDs.
  • Provide at least 75 MB/s of sequential write performance for each serviced HDD.

For write-intensive workloads and iSCSI storage, use NVMe or SAS SSDs. SATA SSDs are generally better suited for SSD-only storage than for cache roles.

On all-SSD and SSD-cached clusters, do not place metadata services on cache SSDs. Use the System+MDS layout instead.

Protecting data during a power outage

To protect Virtuozzo Infrastructure against power outages, it is recommended to use an Uninterruptable Power Supply (UPS) for all servers and network switches.

Additionally, you can prevent data loss and preserve data integrity by using enterprise-grade SSD drives. Unlike HDD drives, enterprise-grade SSD drives can properly handle power loss events. Enterprise-grade SSD drives that operate correctly usually have the power loss protection property in their technical specification. Some of the market names for this technology are Enhanced Power Loss Data Protection (Intel), Cache Power Protection (Samsung), Power-Failure Support (Kingston), and Complete Power Fail Protection (OCZ).

It is also recommended to ensure that all storage devices that will be added to your cluster can flush data from cache to disk in case of a power outage. You can check the data flushing capabilities of your disks by using the procedure below.

Checking disk data flushing capabilities

It is highly recommended to ensure that all storage devices you plan to include in your cluster can flush data from cache to disk if the power goes out unexpectedly. Thus you will find devices that may lose data in a power failure.

Virtuozzo Infrastructure ships with the vstorage-hwflush-check tool that checks how a storage device flushes data to disk in emergencies. The tool is implemented as a client/server utility:

  • The client continuously writes blocks of data to the storage device. When a data block is written, the client increases a special counter and sends it to the server that keeps it.
  • The server keeps track of counters incoming from the client and always knows the next counter number. If the server receives a counter smaller than the one it has (for example, because the power has failed and the storage device has not flushed the cached data to disk), the server reports an error.

To check that a storage device can successfully flush data to disk when power fails, follow the procedure below:

  1. On one node, run the server:
  2. # vstorage-hwflush-check -l
  3. On a different node that hosts the storage device you want to test, run the client. For example:

    # vstorage-hwflush-check -s vstorage1.example.com -d /vstorage/stor1-ssd/test -t 50

    where

    • vstorage1.example.com is the host name of the server.
    • /vstorage/stor1-ssd/test is the directory to use for data flushing tests. During execution, the client creates a file in this directory and writes data blocks to it.
    • 50 is the number of threads for the client to write data to disk. Each thread has its own file and counter. You can increase the number of threads (max. 200) to test your system in more stressful conditions. You can also specify other options when running the client. For more information on available options, refer to the vstorage-hwflush-check manual page.
  4. Wait for at least 10-15 seconds, cut power from the client node (either press the Power button or pull the power cord out), and then power it on again.

  5. Restart the client:
# vstorage-hwflush-check -s vstorage1.example.com -d /vstorlage/stor1-ssd/test -t 50

Once launched, the client will read all previously written data, determine the version of data on the disk, and restart the test from the last valid counter. It then will send this valid counter to the server and the server will compare it to the latest counter it has. You may see output like:

id<N>:<counter_on_disk> -> <counter_on_server>

which means one of the following:

  • If the counter on the disk is lower than the counter on the server, the storage device has failed to flush the data to the disk. Avoid using this storage device in production, especially for CS or journals, as you risk losing data.
  • If the counter on the disk is higher than the counter on the server, the storage device has flushed the data to the disk but the client has failed to report it to the server. The network may be too slow or the storage device may be too fast for the set number of load threads, so consider increasing it. This storage device can be used in production.
  • If both counters are equal, the storage device has flushed the data to the disk and the client has reported it to the server. This storage device can be used in production.

To be on the safe side, repeat the procedure several times. Once you have checked your first storage device, continue with all of the remaining devices you plan to use in the cluster. You need to test all devices you plan to use in the cluster: SSD disks used for CS journaling, disks used for MDS journals and chunk servers.