Language Models are Unsupervised Multitask Learners — GPT-2
GPT-2 amplia la stessa ricetta: più scala e dati fanno emergere capacità zero-shot, anticipando il passaggio dai modelli specializzati ai foundation model.
A living collection of AI experiments, papers I read, and notes that connect the original idea with what stays with me after studying it.
For any questions about these notes, or if you are interested in exchanging ideas and collaborating, feel free to contact me at v.dalfonso@metrica.dev
GPT-2 amplia la stessa ricetta: più scala e dati fanno emergere capacità zero-shot, anticipando il passaggio dai modelli specializzati ai foundation model.
Language Models are Unsupervised Multitask Learners — GPT-2
Dopo il Transformer e in parallelo a BERT, GPT mostra che il pre-training generativo autoregressivo può essere adattato con fine-tuning a numerosi compiti linguistici.
Improving Language Understanding by Generative Pre-Training — GPT-1
Dal Transformer nasce BERT: la self-attention diventa bidirezionale e il pre-training produce rappresentazioni linguistiche riutilizzabili per molti compiti.
BERT: Pre-training of Deep Bidirectional Transformers
Il punto di partenza: sostituisce ricorrenza e convoluzioni con la self-attention, introducendo l’architettura Transformer su cui si costruiscono i paper successivi.
Attention Is All You Need
A living collection of AI experiments, papers I read, and notes that connect the original idea with what stays with me after studying it.
For any questions about these notes, or if you are interested in exchanging ideas and collaborating, feel free to contact me at v.dalfonso@metrica.dev
GPT-2 amplia la stessa ricetta: più scala e dati fanno emergere capacità zero-shot, anticipando il passaggio dai modelli specializzati ai foundation model.
Language Models are Unsupervised Multitask Learners — GPT-2
Dopo il Transformer e in parallelo a BERT, GPT mostra che il pre-training generativo autoregressivo può essere adattato con fine-tuning a numerosi compiti linguistici.
Improving Language Understanding by Generative Pre-Training — GPT-1
Dal Transformer nasce BERT: la self-attention diventa bidirezionale e il pre-training produce rappresentazioni linguistiche riutilizzabili per molti compiti.
BERT: Pre-training of Deep Bidirectional Transformers
Il punto di partenza: sostituisce ricorrenza e convoluzioni con la self-attention, introducendo l’architettura Transformer su cui si costruiscono i paper successivi.
Attention Is All You Need
Colophon
3 topics
A living collection of AI experiments, papers I read, and notes that connect the original idea with what stays with me after studying it.
For any questions about these notes, or if you are interested in exchanging ideas and collaborating, feel free to contact me at v.dalfonso@metrica.dev
GPT-2 amplia la stessa ricetta: più scala e dati fanno emergere capacità zero-shot, anticipando il passaggio dai modelli specializzati ai foundation model.
Language Models are Unsupervised Multitask Learners — GPT-2
Dopo il Transformer e in parallelo a BERT, GPT mostra che il pre-training generativo autoregressivo può essere adattato con fine-tuning a numerosi compiti linguistici.
Improving Language Understanding by Generative Pre-Training — GPT-1
Dal Transformer nasce BERT: la self-attention diventa bidirezionale e il pre-training produce rappresentazioni linguistiche riutilizzabili per molti compiti.
BERT: Pre-training of Deep Bidirectional Transformers
Il punto di partenza: sostituisce ricorrenza e convoluzioni con la self-attention, introducendo l’architettura Transformer su cui si costruiscono i paper successivi.
Attention Is All You Need
Colophon
3 topics
A living collection of AI experiments, papers I read, and notes that connect the original idea with what stays with me after studying it.
For any questions about these notes, or if you are interested in exchanging ideas and collaborating, feel free to contact me at v.dalfonso@metrica.dev
GPT-2 amplia la stessa ricetta: più scala e dati fanno emergere capacità zero-shot, anticipando il passaggio dai modelli specializzati ai foundation model.
Language Models are Unsupervised Multitask Learners — GPT-2
Dopo il Transformer e in parallelo a BERT, GPT mostra che il pre-training generativo autoregressivo può essere adattato con fine-tuning a numerosi compiti linguistici.
Improving Language Understanding by Generative Pre-Training — GPT-1
Dal Transformer nasce BERT: la self-attention diventa bidirezionale e il pre-training produce rappresentazioni linguistiche riutilizzabili per molti compiti.
BERT: Pre-training of Deep Bidirectional Transformers
Il punto di partenza: sostituisce ricorrenza e convoluzioni con la self-attention, introducendo l’architettura Transformer su cui si costruiscono i paper successivi.
Attention Is All You Need
Colophon
3 topics