加载头像
大模型
2024
【论文笔记】Dense Connector for MLLMs
【论文笔记】Dense Connector for MLLMs31
【论文笔记】Attention Prompting on Image for Large  Vision-Language Models
【论文笔记】Attention Prompting on Image for Large Vision-Language Models32
【论文笔记】xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
【论文笔记】xGen-MM (BLIP-3): A Family of Open Large Multimodal Models33
【论文笔记】xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
【论文笔记】xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs34
【论文笔记】X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
【论文笔记】X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs35
【论文笔记】MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding
【论文笔记】MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding36
【论文笔记】Sign2GPT Leveraging Large Language Models for Gloss-Free Sign Language Translation
【论文笔记】Sign2GPT Leveraging Large Language Models for Gloss-Free Sign Language Translation37
【论文笔记】Factorized Learning Assisted with Large Language Model for Gloss-free Sign Language Translation
【论文笔记】Factorized Learning Assisted with Large Language Model for Gloss-free Sign Language Translation38
【论文笔记】VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
【论文笔记】VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs39
【论文笔记】Flamingo: a Visual Language Model for Few-Shot Learning
【论文笔记】Flamingo: a Visual Language Model for Few-Shot Learning40
引用到评论
随便逛逛博客分类文章标签
复制地址关闭热评深色模式轉為繁體