]> www.infradead.org Git - users/hch/misc.git/commitdiff
vfio/type1: optimize vfio_unpin_pages_remote()
authorLi Zhe <lizhe.67@bytedance.com>
Thu, 14 Aug 2025 06:47:14 +0000 (14:47 +0800)
committerAlex Williamson <alex.williamson@redhat.com>
Mon, 6 Oct 2025 17:21:27 +0000 (11:21 -0600)
When vfio_unpin_pages_remote() is called with a range of addresses that
includes large folios, the function currently performs individual
put_pfn() operations for each page. This can lead to significant
performance overheads, especially when dealing with large ranges of pages.

It would be very rare for reserved PFNs and non reserved will to be mixed
within the same range. So this patch utilizes the has_rsvd variable
introduced in the previous patch to determine whether batch put_pfn()
operations can be performed. Moreover, compared to put_pfn(),
unpin_user_page_range_dirty_lock() is capable of handling large folio
scenarios more efficiently.

The performance test results for completing the 16G VFIO IOMMU DMA
unmapping are as follows.

Base(v6.16):
------- AVERAGE (MADV_HUGEPAGE) --------
VFIO UNMAP DMA in 0.141 s (113.7 GB/s)
------- AVERAGE (MAP_POPULATE) --------
VFIO UNMAP DMA in 0.307 s (52.2 GB/s)
------- AVERAGE (HUGETLBFS) --------
VFIO UNMAP DMA in 0.135 s (118.6 GB/s)

With this patchset:
------- AVERAGE (MADV_HUGEPAGE) --------
VFIO UNMAP DMA in 0.044 s (363.2 GB/s)
------- AVERAGE (MAP_POPULATE) --------
VFIO UNMAP DMA in 0.289 s (55.3 GB/s)
------- AVERAGE (HUGETLBFS) --------
VFIO UNMAP DMA in 0.044 s (361.3 GB/s)

For large folio, we achieve an over 67% performance improvement in
the VFIO UNMAP DMA item. For small folios, the performance test
results appear to show a slight improvement.

Suggested-by: Jason Gunthorpe <jgg@ziepe.ca>
Signed-off-by: Li Zhe <lizhe.67@bytedance.com>
Reviewed-by: David Hildenbrand <david@redhat.com>
Acked-by: David Hildenbrand <david@redhat.com>
Link: https://lore.kernel.org/r/20250814064714.56485-6-lizhe.67@bytedance.com
Signed-off-by: Alex Williamson <alex.williamson@redhat.com>
drivers/vfio/vfio_iommu_type1.c

index 30e1b54f6c25f84f0171a39edd10159c9ea05592..916cad80941c457b07de903cd65cc1e5613a443d 100644 (file)
@@ -800,17 +800,29 @@ unpin_out:
        return pinned;
 }
 
+static inline void put_valid_unreserved_pfns(unsigned long start_pfn,
+               unsigned long npage, int prot)
+{
+       unpin_user_page_range_dirty_lock(pfn_to_page(start_pfn), npage,
+                                        prot & IOMMU_WRITE);
+}
+
 static long vfio_unpin_pages_remote(struct vfio_dma *dma, dma_addr_t iova,
                                    unsigned long pfn, unsigned long npage,
                                    bool do_accounting)
 {
        long unlocked = 0, locked = vpfn_pages(dma, iova, npage);
-       long i;
 
-       for (i = 0; i < npage; i++)
-               if (put_pfn(pfn++, dma->prot))
-                       unlocked++;
+       if (dma->has_rsvd) {
+               unsigned long i;
 
+               for (i = 0; i < npage; i++)
+                       if (put_pfn(pfn++, dma->prot))
+                               unlocked++;
+       } else {
+               put_valid_unreserved_pfns(pfn, npage, dma->prot);
+               unlocked = npage;
+       }
        if (do_accounting)
                vfio_lock_acct(dma, locked - unlocked, true);