References
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. (2021). On the opportunities and risks of foundation models. arXiv:2108.07258. π Open
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877β1901. π Open
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., & Fei-Fei, L. (2009). ImageNet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition (pp. 248β255). π Open
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. π Open
Fukushima, K. (1980). Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position. Biological Cybernetics, 36(4), 193β202. π Open
Giesen, J., Nussbaum, F., & Schneider, C. (2019). Efficient regularization parameter selection for latent variable graphical models via bi-level optimization. In Proceedings of the 28th International Joint Conference on Artificial Intelligence (pp. 2378β2384). AAAI Press.
Giesen, J., Kahlmeyer, P., Laue, S., Mitterreiter, M., Nussbaum, F., Staudt, C., & ZarrieΓ, S. (2021). Method of moments for topic models with mixed discrete and continuous features. In Proceedings of the 30th International Joint Conference on Artificial Intelligence (pp. 2418β2424).
Giesen, J., Kahlmeyer, P., Nussbaum, F., & ZarrieΓ, S. (2022). Leveraging the Wikipedia graph for evaluating word embeddings. In Proceedings of the 31st International Joint Conference on Artificial Intelligence (pp. 4136β4142).
Giesen, J., Kahlmeyer, P., Laue, S., Mitterreiter, M., Nussbaum, F., & Staudt, C. (2023). Mixed membership Gaussians. Journal of Multivariate Analysis, 195, 105141.
He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 770β778). π Open
Hinton, G. E., & Salakhutdinov, R. R. (2006). Reducing the dimensionality of data with neural networks. Science, 313(5786), 504β507. π Open
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., & Amodei, D. (2020). Scaling laws for neural language models. arXiv:2001.08361. π Open
Kingma, D. P., & Ba, J. (2015). Adam: A method for stochastic optimization. International Conference on Learning Representations. π Open
LeCun, Y., Bottou, L., Bengio, Y., & Haffner, P. (1998). Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11), 2278β2324. π Open
LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436β444. π Open
Nussbaum, F., & Giesen, J. (2019). Ising models with latent conditional Gaussian variables. In Proceedings of the 30th International Conference on Algorithmic Learning Theory (vol. 98, pp. 669β681). PMLR.
Nussbaum, F., & Giesen, J. (2020). Disentangling direct and indirect interactions in polytomous item response theory models. In Proceedings of the 29th International Joint Conference on Artificial Intelligence (IJCAI-20) (pp. 2241β2247). International Joint Conferences on Artificial Intelligence Organization. π Open
Nussbaum, F., & Giesen, J. (2020). Pairwise sparse + low-rank models for variables of mixed type. Journal of Multivariate Analysis, 178, 104601. π Open
Nussbaum, F. (2021). Models with low-rank and group-sparse components and their recovery via convex optimization [Doctoral dissertation].
Nussbaum, F., & Giesen, J. (2021). Robust principal component analysis for generalized multi-view models. In Uncertainty in Artificial Intelligence (pp. 686β695). PMLR.
Nussbaum, F., Gawlikowski, J., & Niebling, J. (2022). Structuring uncertainty for fine-grained sampling in stochastic segmentation networks. Advances in Neural Information Processing Systems, 35, 27678β27691.
Nussbaum, F. G. (2023). Comprehensive review of AI myths and misconceptions.
Nussbaum, F. G. (2023). Successful communication of complex information. ResearchGate Preprint.
Nussbaum, F. G. (2025). Plan. Pitch. Perform. From data science idea to funded project.
Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018). Improving language understanding by generative pre-training. OpenAI Technical Report. π Open
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., & Liu, P. J. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140), 1β67. π Open
Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323(6088), 533β536. π Open
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., & Salakhutdinov, R. (2014). Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15(56), 1929β1958. π Open
Tan, P.-N., Steinbach, M., Karpatne, A., & Kumar, V. (2020). Introduction to Data Mining (2nd ed.). Pearson.
Van der Maaten, L., & Hinton, G. (2008). Visualizing data using t-SNE. Journal of Machine Learning Research, 9(86), 2579β2605. π Open
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30. π Open
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., & Fedus, W. (2022). Emergent abilities of large language models. Transactions on Machine Learning Research. π Open
Yosinski, J., Clune, J., Bengio, Y., & Lipson, H. (2014). How transferable are features in deep neural networks? Advances in Neural Information Processing Systems, 27, 3320β3328. π Open