加载头像
论文笔记
2024
【论文笔记】VCoder: Versatile Vision Encoders for Multimodal Large Language Models
【论文笔记】VCoder: Versatile Vision Encoders for Multimodal Large Language Models11
【论文笔记】Dense Connector for MLLMs
【论文笔记】Dense Connector for MLLMs12
【论文笔记】Attention Prompting on Image for Large  Vision-Language Models
【论文笔记】Attention Prompting on Image for Large Vision-Language Models13
【论文笔记】Token Turing Machines
【论文笔记】Token Turing Machines14
【论文笔记】Gloss-free Sign Language Translation: Improving from Visual-Language Pretraining
【论文笔记】Gloss-free Sign Language Translation: Improving from Visual-Language Pretraining15
【论文笔记】C$^2$RL: Content and Context Representation Learning for Gloss-free Sign Language Translation and Retrieval
【论文笔记】C$^2$RL: Content and Context Representation Learning for Gloss-free Sign Language Translation and Retrieval16
【论文笔记】Perceiver: General Perception with Iterative Attention
【论文笔记】Perceiver: General Perception with Iterative Attention17
【论文笔记】xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
【论文笔记】xGen-MM (BLIP-3): A Family of Open Large Multimodal Models18
【论文笔记】xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
【论文笔记】xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs19
【论文笔记】MLSLT: Towards Multilingual Sign Language Translation
【论文笔记】MLSLT: Towards Multilingual Sign Language Translation20
引用到评论
随便逛逛博客分类文章标签
复制地址关闭热评深色模式轉為繁體