Smooth Attention: Improving Image Semantic Segmentation
DOI:
https://doi.org/10.34257/LJRCSTVol24IS2PG17Keywords:
design software, graphic design., OTN (optical transport network), optical control plane, SDON (software defined optical network), optiSystem, openflow, SDN (software defined network)., photonics., optics, light, lasers, journal manuscripts, LaTeX template.Abstract
This document shows the required format and appearance of a manuscript prepared for SPIE journals.
It is prepared using LaTeX2e with the class file spieman.cls. Please note that the following journals require the use of structured abstracts in manuscript submissions: Neurophotonics, the Journal of Biomedical Optics, and the Journal of Medical Imaging. Structured abstracts are encouraged for the Journal of Micro/Nanolithography, MEMS, and MOEMS. Guidelines are available on the journal website. Whether structured or single-paragraph, the abstract should be a summary of the paper and not an introduction. Because the abstract may be used in abstracting and indexing databases, it should be self-contained (i.e., no numerical references) and substantive in nature, presenting concisely the objectives, methodology used, results obtained, and their significance. A list of up to six keywords should immediately follow.
References
A. Vaswani (2017) Attention is all you need.
A. Galassi, M. Lippi, P. Torroni (2020) Attention in natural language processing. 32(10), 4291-4308.
M. H. Guo, T. X. Xu, J. J. Liu, Z. N. Liu, P. T. Jiang, T. J. Mu, S. M. Hu (2022) Attention mechanisms in computer vision: A survey. 8(3), 331-368.
X. Yang (2020) An overview of the attention mechanisms in computer vision. 1693(1), 012173.
F. Wang, D. M. Tax (2016) Survey on the attention based RNN model and its applications in computer vision. https://arxiv.org/abs/1601.06823
I. Bello, B. Zoph, A. Vaswani, J. Shlens, Q. V. Le (2019) Attention augmented convolutional networks. 3286-3295.
P. Ramachandran, N. Parmar, A. Vaswani, I. Bello, A. Levskaya, J. Shlens (2019) Stand-alone self-attention in vision models.
M. T. Luong (2015) Effective approaches to attention-based neural machine translation. https://arxiv.org/abs/1508.04025
B. Zhang, D. Xiong, J. Su (2018) Neural machine translation with deep attention.
S. Maruf, A. F. Martins, G. Haffari (2019) Selective attention for context-aware neural machine translation. https://arxiv.org/abs/1903.08788
Q. You, H. Jin, Z. Wang, C. Fang, J. Luo (2016) Image captioning with semantic attention. 4651-4659.
J. Sun, J. Jiang, Y. Liu (2020) An introductory survey on attention mechanisms in computer vision problems. 295-300.
H. Li, P. Xiong, J. An, L. Wang (2018) Pyramid attention network for semantic segmentation. https://arxiv.org/abs/1805.10180
M. H. Guo, C. Z. Lu, Q. Hou, Z. Liu, M. M. Cheng, S. M. Hu (2022) Segnext: Rethinking convolutional attention design for semantic segmentation. 35, 1140-1156.
S. Konate, L. Lebrat, R. Santa Cruz, P. Bourgeat, V. Dore, J. Fripp, O. Salvado (2021) Smocam: Smooth conditional attention mask for 3d-regression models. 362-366.
Y. Yao, J. Ren, X. Xie, W. Liu, Y. J. Liu, J. Wang (2019) Attention-aware multi-stroke style transfer. 1467-1475.
P. T. Jiang, L. H. Han, Q. Hou, M. M. Cheng, Y. Wei (2021) Online attention accumulation for weakly supervised semantic segmentation. 44(10), 7062-7077.
Alexey D. (2020) An image is worth 16x16 words: Transformers for image recognition at scale. https://arxiv.org/abs/2010.11929
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, B. Guo (2021) Swin transformer: Hierarchical vision transformer using shifted windows. 10012-10022.
X. Wang, R. Girshick, A. Gupta, K. He (2018) Non-local neural networks. 7794-7803.
S. Woo, J. Park, J. Y. Lee, I. S. Kweon (2018) Cbam: Convolutional block attention module. 3-19.
S. Zheng, S. Jayasumana, B. Romera-Paredes, V. Vineet, Z. Su, D. Du, P. H. Torr (2015) Conditional random fields as recurrent neural networks. 1529-1537.
P. Isola, J. Y. Zhu, T. Zhou, A. A. Efros (2017) Image-to-image translation with conditional adversarial networks. 1125-1134.
A. Graves (2016) Adaptive computation time for recurrent neural networks. https://arxiv.org/abs/1603.08983
J. Yang, C. Li, P. Zhang, X. Dai, B. Xiao, L. Yuan, J. Gao (2021) Focal self-attention for local-global interactions in vision transformers. https://arxiv.org/abs/2107.00641
J. Wang, Y. Chen, S. Hao, X. Peng, L. Hu (2019) Deep learning for sensor-based activity recognition: A survey. 119, 3-11.
R. Liu, J. Lehman, P. Molino, F. Petroski Such, E. Frank, A. Sergeev, J. Yosinski (2018) An intriguing failing of convolutional neural networks and the coordconv solution.
K. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, A. Weller (2020) Rethinking attention with performers. https://arxiv.org/abs/2009.14794
J. Solomon, K. Crane, A. Butscher, C. Wojtan (2014) A general framework for bilateral and mean shift filtering. 1(2), 3. https://arxiv.org/abs/1405.4734
J. Wang, X. Hu (2021) Convolutional neural networks with gated recurrent connections. 44(7), 3421-3435.
S. Volz, A. Bruhn, L. Valgaerts, H. Zimmer (2011) Modeling temporal coherence for optical flow. 1116-1123.
X. Tong, R. Xu, K. Liu, L. Zhao, W. Zhu, D. Zhao (2023) A Deep-Learning Approach for Low-Spatial-Coherence Imaging in Computer-Generated Holography. 4(1), 2200264.
R. Zabih, V. Kolmogorov (2004) Spatially coherent clustering using graph cuts. 2, II-II.
F. Li, G. Lebanon, C. Sminchisescu (2012) Chebyshev approximations to the histogram X 2 kernel. 2424-2431.
J. Koenderink, A. van Doom (1998) Shape from Chebyshev nets. 5, 215-225.
C. Wah, S. Branson, P. Welinder, P. Perona, S. Belongie (2011) The caltech-ucsd birds-200-2011 dataset.
O. Ulucan, D. Karakaya, M. Turkan (2020) A large-scale dataset for fish segmentation and classification. 1-5.
DiversisAI Fire segmentation image dataset. https://www.kaggle.com/datasets/diversisai/fire-segmentation-image-dataset
D. Jha, S. Ali, K. Emanuelsen, S. A. Hicks, V. Thambawita, E. Garcia-Ceja, M. A. Riegler, T. de Lange, P. T. Schmidt, H. D. Johansen, D. Johansen, P. Halvorsen (2021) Kvasir-Instrument: Diagnostic and therapeutic tool segmentation dataset in gastrointestinal endoscopy. 218-229.
H. Li Flood semantic segmentation dataset. https://www.kaggle.com/datasets/lihuayang111265/flood-semantic-segmentation-dataset
N. Siddique, S. Paheding, C. P. Elkin, V. Devabhaktuni (2021) U-net and its variants for medical image segmentation: A review of theory and applications. 9, 82031-82057.
O. Ronneberger, P. Fischer, T. Brox (2015) U-net: Convolutional networks for biomedical image segmentation. 18, 234-241.
O. Oktay, J. Schlemper, L. L. Folgoc, M. Lee, M. Heinrich, K. Misawa, D. Rueckert (2018) Attention u-net: Learning where to look for the pancreas. http://arxiv.org/abs/1804.03999
K. He, X. Zhang, S. Ren, J. Sun (2016) Deep residual learning for image recognition. 770-778.
H. Gao, H. Yuan, Z. Wang, S. Ji (2019) Pixel transposed convolutional networks. 42(5), 1218-1227.
K. Cho, A. Courville, Y. Bengio (2015) Describing multimedia content using attention-based encoder-decoder networks. 17(11), 1875-1886.
Z. Ji, K. Xiong, Y. Pang, X. Li (2019) Video summarization with attention-based encoder-decoder networks. 30(6), 1709-1717.
A. Garcia-Garcia, S. Orts-Escolano, S. Oprea, V. Villena-Martinez, J. Garcia-Rodriguez (2017) A review on deep learning techniques applied to semantic segmentation. http://arxiv.org/abs/1704.06857
S. Du, H. Fan, M. Zhao, H. Zong, J. Hu, P. Li (2022) A two-stage method for single image de-raining based on attention smoothed dilated networks. 16(10), 2557-2567.
W. Ouyang, X. Zeng, X. Wang (2013) Modeling mutual visibility relationship in pedestrian detection. 3222-3229.
Downloads
Published
Issue
Section
License
Copyright (c) 2024 Authors and Global Journals Private Limited

This work is licensed under a Creative Commons Attribution 4.0 International License.
