Skip to content

References

Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. (2021). On the opportunities and risks of foundation models. arXiv:2108.07258. πŸ”— Open

Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901. πŸ”— Open

Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., & Fei-Fei, L. (2009). ImageNet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition (pp. 248–255). πŸ”— Open

Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. πŸ”— Open

Fukushima, K. (1980). Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position. Biological Cybernetics, 36(4), 193–202. πŸ”— Open

Giesen, J., Nussbaum, F., & Schneider, C. (2019). Efficient regularization parameter selection for latent variable graphical models via bi-level optimization. In Proceedings of the 28th International Joint Conference on Artificial Intelligence (pp. 2378–2384). AAAI Press.

Giesen, J., Kahlmeyer, P., Laue, S., Mitterreiter, M., Nussbaum, F., Staudt, C., & Zarrieß, S. (2021). Method of moments for topic models with mixed discrete and continuous features. In Proceedings of the 30th International Joint Conference on Artificial Intelligence (pp. 2418–2424).

Giesen, J., Kahlmeyer, P., Nussbaum, F., & Zarrieß, S. (2022). Leveraging the Wikipedia graph for evaluating word embeddings. In Proceedings of the 31st International Joint Conference on Artificial Intelligence (pp. 4136–4142).

Giesen, J., Kahlmeyer, P., Laue, S., Mitterreiter, M., Nussbaum, F., & Staudt, C. (2023). Mixed membership Gaussians. Journal of Multivariate Analysis, 195, 105141.

He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 770–778). πŸ”— Open

Hinton, G. E., & Salakhutdinov, R. R. (2006). Reducing the dimensionality of data with neural networks. Science, 313(5786), 504–507. πŸ”— Open

Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., & Amodei, D. (2020). Scaling laws for neural language models. arXiv:2001.08361. πŸ”— Open

Kingma, D. P., & Ba, J. (2015). Adam: A method for stochastic optimization. International Conference on Learning Representations. πŸ”— Open

LeCun, Y., Bottou, L., Bengio, Y., & Haffner, P. (1998). Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11), 2278–2324. πŸ”— Open

LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436–444. πŸ”— Open

Nussbaum, F., & Giesen, J. (2019). Ising models with latent conditional Gaussian variables. In Proceedings of the 30th International Conference on Algorithmic Learning Theory (vol. 98, pp. 669–681). PMLR.

Nussbaum, F., & Giesen, J. (2020). Disentangling direct and indirect interactions in polytomous item response theory models. In Proceedings of the 29th International Joint Conference on Artificial Intelligence (IJCAI-20) (pp. 2241–2247). International Joint Conferences on Artificial Intelligence Organization. πŸ”— Open

Nussbaum, F., & Giesen, J. (2020). Pairwise sparse + low-rank models for variables of mixed type. Journal of Multivariate Analysis, 178, 104601. πŸ”— Open

Nussbaum, F. (2021). Models with low-rank and group-sparse components and their recovery via convex optimization [Doctoral dissertation].

Nussbaum, F., & Giesen, J. (2021). Robust principal component analysis for generalized multi-view models. In Uncertainty in Artificial Intelligence (pp. 686–695). PMLR.

Nussbaum, F., Gawlikowski, J., & Niebling, J. (2022). Structuring uncertainty for fine-grained sampling in stochastic segmentation networks. Advances in Neural Information Processing Systems, 35, 27678–27691.

Nussbaum, F. G. (2023). Comprehensive review of AI myths and misconceptions.

Nussbaum, F. G. (2023). Successful communication of complex information. ResearchGate Preprint.

Nussbaum, F. G. (2025). Plan. Pitch. Perform. From data science idea to funded project.

Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018). Improving language understanding by generative pre-training. OpenAI Technical Report. πŸ”— Open

Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., & Liu, P. J. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140), 1–67. πŸ”— Open

Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323(6088), 533–536. πŸ”— Open

Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., & Salakhutdinov, R. (2014). Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15(56), 1929–1958. πŸ”— Open

Tan, P.-N., Steinbach, M., Karpatne, A., & Kumar, V. (2020). Introduction to Data Mining (2nd ed.). Pearson.

Van der Maaten, L., & Hinton, G. (2008). Visualizing data using t-SNE. Journal of Machine Learning Research, 9(86), 2579–2605. πŸ”— Open

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30. πŸ”— Open

Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., & Fedus, W. (2022). Emergent abilities of large language models. Transactions on Machine Learning Research. πŸ”— Open

Yosinski, J., Clune, J., Bengio, Y., & Lipson, H. (2014). How transferable are features in deep neural networks? Advances in Neural Information Processing Systems, 27, 3320–3328. πŸ”— Open