In the context of VPR, we want to match a query image against references from an extensive database, relying solely on visual cues. State-of-the-art pipelines focus on the aggregation of features extracted from a deep backbone, in order to form a global descriptor for each image.
SALAD
Itβs a feature aggregation module.
They consider both feature-to-cluster and cluster-to-feature relations and introduce a βdustbinβ cluster, designed to selectively discard features deemed non-informative.
![]()
For my application regarding VPR, it could act as a filter against partial correspondences and task-irrelevant background regions (sky).
Since DINOv2 and VGGT both use patch size, the confidence maps would align 1:1. The distribution will differ massively, however, since VGGTβs features encode structural constraints and spatial relationships rather than semantic meaning (DINO). However, I expect the resulting confidence maps to highlight structural correspondences rather than semantic objects.
CLIP ViT-B/16 uses patches, so I would have to interpolate.
What I want is the optimal transport side of it. I donβt care about the fine-tuned DINOv2 backbone.
Reduce assignment priors
In NetVLAD, a global descriptor is formed by assigning a set of features to a set of clusters . Then, NetVLAD computes a score matrix where each element is the cost of assigning a feature to a cluster .
Priorly, was initialized with centroids derived from k-means. It accelerated training but it introduced bias and made the model more susceptible to local minima. SALAD learns each row from scratch with two fully connected layers initialized randomly:
Discard uninformative features
Additionally they augment by adding a column representing the feature-to-dustbin relation. This score is modeled with a single learnable parameter :
Optimal assignment
NetVLAD computes a per-row softmax over to obtain the distribution of each featureβs mass across the clusters. Since this approach overlooks the custer-to-feature relation, SALAD reformulates it as an optimal transport problem where the featuresβ mass, , must be effectively distributed among the clusters or the dustbin, .
They use the Sinkhorn Algorithm to obtain the assignment such that:
Finally, they drop the dustbin column to obtain the assignment .
Dimensionality reduction
To manage the final descriptor size, they reduce the dimensionality of the tokens from to . This is again achieved by processing the features through two fully connected layers which adjust the size of the feature vectors while retaining essential information from the task.
Aggregation
Differently from NetVLAD, they do not subtract the centroids to get the residuals, but directly aggregate the features with a summation, reducing the incorporated priors about the aggregation. It results in the following VLAD vector as a matrix . Each element is computed as follows:
Global token
To include global information about the scene (which is not easily incorporated into local features), they also incorporate a scene descriptor framed exactly as other two fully connected layers:
- is the global token from DINOv2.
What do I use for VGGT for ?
In the end they concatenate with flattened, followed by an L2 intra-normalization and an entire L2 normalization of this vector final global descriptor.