A living collection of AI experiments, papers I read, and notes that connect the original idea with what stays with me after studying it.

For any questions about these notes, or if you are interested in exchanging ideas and collaborating, feel free to contact me at v.dalfonso@metrica.dev

The Journey04

Language Models are Unsupervised Multitask Learners — GPT-2

GPT-2 amplia la stessa ricetta: più scala e dati fanno emergere capacità zero-shot, anticipando il passaggio dai modelli specializzati ai foundation model.

Language Models are Unsupervised Multitask Learners — GPT-2

The Journey03

Improving Language Understanding by Generative Pre-Training — GPT-1

Dopo il Transformer e in parallelo a BERT, GPT mostra che il pre-training generativo autoregressivo può essere adattato con fine-tuning a numerosi compiti linguistici.

Improving Language Understanding by Generative Pre-Training — GPT-1

The Journey02

BERT: Pre-training of Deep Bidirectional Transformers

Dal Transformer nasce BERT: la self-attention diventa bidirezionale e il pre-training produce rappresentazioni linguistiche riutilizzabili per molti compiti.

BERT: Pre-training of Deep Bidirectional Transformers

The Journey01

Attention Is All You Need

Il punto di partenza: sostituisce ricorrenza e convoluzioni con la self-attention, introducendo l’architettura Transformer su cui si costruiscono i paper successivi.

Attention Is All You Need

A living collection of AI experiments, papers I read, and notes that connect the original idea with what stays with me after studying it.

For any questions about these notes, or if you are interested in exchanging ideas and collaborating, feel free to contact me at v.dalfonso@metrica.dev

The Journey04

Language Models are Unsupervised Multitask Learners — GPT-2

GPT-2 amplia la stessa ricetta: più scala e dati fanno emergere capacità zero-shot, anticipando il passaggio dai modelli specializzati ai foundation model.

Language Models are Unsupervised Multitask Learners — GPT-2

The Journey03

Improving Language Understanding by Generative Pre-Training — GPT-1

Dopo il Transformer e in parallelo a BERT, GPT mostra che il pre-training generativo autoregressivo può essere adattato con fine-tuning a numerosi compiti linguistici.

Improving Language Understanding by Generative Pre-Training — GPT-1

The Journey02

BERT: Pre-training of Deep Bidirectional Transformers

Dal Transformer nasce BERT: la self-attention diventa bidirezionale e il pre-training produce rappresentazioni linguistiche riutilizzabili per molti compiti.

BERT: Pre-training of Deep Bidirectional Transformers

The Journey01

Attention Is All You Need

Il punto di partenza: sostituisce ricorrenza e convoluzioni con la self-attention, introducendo l’architettura Transformer su cui si costruiscono i paper successivi.

Attention Is All You Need

A living collection of AI experiments, papers I read, and notes that connect the original idea with what stays with me after studying it.

For any questions about these notes, or if you are interested in exchanging ideas and collaborating, feel free to contact me at v.dalfonso@metrica.dev

The Journey04

Language Models are Unsupervised Multitask Learners — GPT-2

GPT-2 amplia la stessa ricetta: più scala e dati fanno emergere capacità zero-shot, anticipando il passaggio dai modelli specializzati ai foundation model.

Language Models are Unsupervised Multitask Learners — GPT-2

The Journey03

Improving Language Understanding by Generative Pre-Training — GPT-1

Dopo il Transformer e in parallelo a BERT, GPT mostra che il pre-training generativo autoregressivo può essere adattato con fine-tuning a numerosi compiti linguistici.

Improving Language Understanding by Generative Pre-Training — GPT-1

The Journey02

BERT: Pre-training of Deep Bidirectional Transformers

Dal Transformer nasce BERT: la self-attention diventa bidirezionale e il pre-training produce rappresentazioni linguistiche riutilizzabili per molti compiti.

BERT: Pre-training of Deep Bidirectional Transformers

The Journey01

Attention Is All You Need

Il punto di partenza: sostituisce ricorrenza e convoluzioni con la self-attention, introducendo l’architettura Transformer su cui si costruiscono i paper successivi.

Attention Is All You Need

A living collection of AI experiments, papers I read, and notes that connect the original idea with what stays with me after studying it.

For any questions about these notes, or if you are interested in exchanging ideas and collaborating, feel free to contact me at v.dalfonso@metrica.dev

The Journey04

Language Models are Unsupervised Multitask Learners — GPT-2

GPT-2 amplia la stessa ricetta: più scala e dati fanno emergere capacità zero-shot, anticipando il passaggio dai modelli specializzati ai foundation model.

Language Models are Unsupervised Multitask Learners — GPT-2

The Journey03

Improving Language Understanding by Generative Pre-Training — GPT-1

Dopo il Transformer e in parallelo a BERT, GPT mostra che il pre-training generativo autoregressivo può essere adattato con fine-tuning a numerosi compiti linguistici.

Improving Language Understanding by Generative Pre-Training — GPT-1

The Journey02

BERT: Pre-training of Deep Bidirectional Transformers

Dal Transformer nasce BERT: la self-attention diventa bidirezionale e il pre-training produce rappresentazioni linguistiche riutilizzabili per molti compiti.

BERT: Pre-training of Deep Bidirectional Transformers

The Journey01

Attention Is All You Need

Il punto di partenza: sostituisce ricorrenza e convoluzioni con la self-attention, introducendo l’architettura Transformer su cui si costruiscono i paper successivi.

Attention Is All You Need