arXiv:2407.01782 Abstract | arXiv Analytics

arXiv:2407.01782 [cs.CV]Abstract References Reviews Resources

Addressing a fundamental limitation in deep vision models: lack of spatial attention

Published 2024-07-01Version 1

The primary aim of this manuscript is to underscore a significant limitation in current deep learning models, particularly vision models. Unlike human vision, which efficiently selects only the essential visual areas for further processing, leading to high speed and low energy consumption, deep vision models process the entire image. In this work, we examine this issue from a broader perspective and propose a solution that could pave the way for the next generation of more efficient vision models. Basically, convolution and pooling operations are selectively applied to altered regions, with a change map sent to subsequent layers. This map indicates which computations need to be repeated. The code is available at https://github.com/aliborji/spatial_attention.

Categories: cs.CV, cs.AI

Keywords: spatial attention, fundamental limitation, deep vision models process, essential visual areas, current deep learning models

Related articles: Most relevant | Search more

arXiv:2307.07370 [cs.CV] (Published 2023-07-14)

AIC-AB NET: A Neural Network for Image Captioning with Spatial Attention and Text Attributes

Guoyun Tu, Ying Liu, Vladimir Vlassov

arXiv:2401.02656 [cs.CV] (Published 2024-01-05)

GTA: Guided Transfer of Spatial Attention from Object-Centric Representations

SeokHyun Seo, Jinwoo Hong, JungWoo Chae, Kyungyul Kim, Sangheum Hwang

arXiv:2104.09807 [cs.CV] (Published 2021-04-20)

Visual Navigation with Spatial Attention

Bar Mayo, Tamir Hazan, Ayellet Tal