Re-se-arch – The Laboratory for interpretable Visual Modeling, Computing and Learning (iVMCL)

Grainger, Ryan; Paniagua, Thomas; Song, Xi; Wu, Tianfu

Learning Patch-to-Cluster Attention in Vision Transformer Working paper

arXiv preprint, 2022.

Abstract | Links | BibTeX

Xue, Nan; Wu, Tianfu; Bai, Song; Wang, Fudong; Xia, Gui-Song; Zhang, Liangpei; Torr, Philip H. S.

Holistically-Attracted Wireframe Parsing Proceedings Article

In: IEEE Conference on Computer Vision and Pattern Recognition (CVRP), 2020., 2020.

Abstract | BibTeX

Xing, Xianglei; Wu, Tianfu; Zhu, Song-Chun; Wu, Ying Nian

Towards Interpretable Image Synthesis by Learning Sparsely Connected AND-OR Networks Proceedings Article

In: IEEE Conference on Computer Vision and Pattern Recognition (CVRP), 2020., 2020.

Abstract | Links | BibTeX

@inproceedings{iGenerativeM,

title = {Towards Interpretable Image Synthesis by Learning Sparsely Connected AND-OR Networks},

author = {Xianglei Xing and Tianfu Wu and Song-Chun Zhu and Ying Nian Wu},

url = {https://arxiv.org/abs/1909.04324},

year  = {2020},

date = {2020-02-23},

booktitle = {IEEE Conference on Computer Vision and Pattern Recognition (CVRP), 2020.},

journal = {CoRR},

abstract = {This paper proposes interpretable image synthesis by learning hierarchical AND-OR networks of sparsely connected semantically meaningful nodes. The proposed method is based on the compositionality and interpretability of scene-objects-parts-subparts-primitives hierarchy in image representation. A scene has different types (i.e., OR) each of which consists of a number of objects (i.e., AND). This can be recursively formulated across the scene-objects-parts-subparts hierarchy and is terminated at the primitive level (e.g., Gabor wavelets-like basis). To realize this interpretable AND-OR hierarchy in image synthesis, the proposed method consists of two components: (i) Each layer of the hierarchy is represented by an over-completed set of basis functions. The basis functions are instantiated using convolution to be translation covariant. Off-the-shelf convolutional neural architectures are then exploited to implement the hierarchy. (ii) Sparsity-inducing constraints are introduced in end-to-end training, which facilitate a sparsely connected AND-OR network to emerge from initially densely connected convolutional neural networks. A straightforward sparsity-inducing constraint is utilized, that is to only allow the top-k basis functions to be active at each layer (where k is a hyperparameter). The learned basis functions are also capable of image reconstruction to explain away input images. In experiments, the proposed method is tested on five benchmark datasets. The results show that meaningful and interpretable hierarchical representations are learned with better qualities of image synthesis and reconstruction obtained than state-of-the-art baselines.},

howpublished = {IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020},

keywords = {},

pubstate = {published},

tppubtype = {inproceedings}

}

Close

Li, Xilai; Song, Xi; Wu, Tianfu

AOGNets: Compositional Grammatical Architectures for Deep Learning Proceedings Article

In: IEEE Conference on Computer Vision and Pattern Recognition (CVRP), 2019.

Abstract | Links | BibTeX

@inproceedings{AOGNets,

title = {AOGNets: Compositional Grammatical Architectures for Deep Learning},

author = {Xilai Li and Xi Song and Tianfu Wu},

url = {http://openaccess.thecvf.com/content_CVPR_2019/papers/Li_AOGNets_Compositional_Grammatical_Architectures_for_Deep_Learning_CVPR_2019_paper.pdf 

https://github.com/iVMCL/AOGNets

https://www.wraltechwire.com/2019/05/21/ncsu-researchers-create-framework-for-a-smarter-ai-are-seeking-patent/

https://www.technologynetworks.com/tn/news/new-framework-enhances-neural-network-performance-319704

},

year  = {2019},

date = {2019-06-18},

booktitle = {IEEE Conference on Computer Vision and Pattern Recognition (CVRP)},

abstract = {Neural architectures are the foundation for improving performance of deep neural networks (DNNs). This paper presents deep compositional grammatical architectures which harness the best of two worlds: grammar models and DNNs. The proposed architectures integrate compositionality and reconfigurability of the former and the capability of learning rich features of the latter in a principled way. We utilize AND-OR Grammar (AOG) as network generator in this paper and call the resulting networks AOGNets. An AOGNet consists of a number of stages each of which is composed of a number of AOG building blocks. An AOG building block splits its input feature map into N groups along feature channels and then treat it as a sentence of N words. It then jointly realizes a phrase structure grammar and a dependency grammar in bottom-up parsing the “sentence” for better feature exploration and reuse. It provides a unified framework for the best practices developed in state-of-the-art DNNs. In experiments, AOGNet is tested in the ImageNet-1K classification benchmark and the MS-COCO object detection and segmentation benchmark. In ImageNet-1K, AOGNet obtains better performance than ResNet and most of its variants, ResNeXt and its attention based variants such as SENet, DenseNet and DualPathNet. AOGNet also obtains the best model interpretability score using network dissection. AOGNet further shows better potential in adversarial defense. In MS-COCO, AOGNet obtains better performance than the ResNet and ResNeXt backbones in Mask R-CNN.},

keywords = {},

pubstate = {published},

tppubtype = {inproceedings}

}

Close

Sun, Wei; Wu, Tianfu

Image Synthesis from Reconfigurable Layout and Style Proceedings Article

In: International Conference on Computer Vision (ICCV), 2019.

Abstract | Links | BibTeX

Li, Xilai; Zhou, Yingbo; Wu, Tianfu; Socher, Richard; Xiong, Caiming

Learn to Grow: A Continual Structure Learning Framework for Overcoming Catastrophic Forgetting Proceedings Article

In: International Conference on Machine Learning (ICML), 2019.

Abstract | Links | BibTeX

Wu, Tianfu; Song, Xi

Towards Interpretable Object Detection by Unfolding Latent Structures Proceedings Article

In: International Conference on Computer Vision (ICCV), 2019.

Abstract | BibTeX