About calculation of loss function #28

Littleor · 2021-11-27T11:35:52Z

Hi, I have some questions about this loss function because there are some differences between the paper and the code.
First of all, the paper:

The loss is the sum of the average logsoftmax between the query point and other prototypes and the average distance between the query point and the corresponding prototypes.

But in the code:

       log_p_y = F.log_softmax(-dists, dim=1).view(n_class, n_query, -1)
       loss_val = -log_p_y.gather(2, target_inds).squeeze().view(-1).mean()

The loss in code is only the sum of LogSoftmax between the query point and the corresponding prototypes.

I'm confused. Is it my understanding of the code or my understanding of the paper?

fabian57fabian · 2023-09-29T12:46:43Z

I had the same confusion when replicated the paper's results.
You could basically run cross-entropy and results will be the same.
In this formula, he uses log_softmax on distances, then take only the right ones with gather function and compute mean on that.
So it is basically the same.

I did similarly in my own implementation:
https://github.com/fabian57fabian/prototypical-networks-few-shot-learning

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

About calculation of loss function #28

About calculation of loss function #28

Littleor commented Nov 27, 2021 •

edited

Loading

fabian57fabian commented Sep 29, 2023

About calculation of loss function #28

About calculation of loss function #28

Comments

Littleor commented Nov 27, 2021 • edited Loading

fabian57fabian commented Sep 29, 2023

Littleor commented Nov 27, 2021 •

edited

Loading