Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Speech is inherently easier to represent as a sequence of tokens than a high-resolution image.

Best speech to text is already NN transformer based anyway, so in theory it's only better to use a combined model



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: