Impact of Noisy Supervision in Foundation Model Learning

Chen, Hao; Wang, Zihan; Tao, Ran; Wei, Hongxin; Xie, Xing; Sugiyama, Masashi; Raj, Bhiksha; Wang, Jindong

Computer Science > Machine Learning

arXiv:2403.06869 (cs)

[Submitted on 11 Mar 2024 (v1), last revised 5 May 2025 (this version, v3)]

Title:Impact of Noisy Supervision in Foundation Model Learning

Authors:Hao Chen, Zihan Wang, Ran Tao, Hongxin Wei, Xing Xie, Masashi Sugiyama, Bhiksha Raj, Jindong Wang

View PDF HTML (experimental)

Abstract:Foundation models are usually pre-trained on large-scale datasets and then adapted to downstream tasks through tuning. However, the large-scale pre-training datasets, often inaccessible or too expensive to handle, can contain label noise that may adversely affect the generalization of the model and pose unexpected risks. This paper stands out as the first work to comprehensively understand and analyze the nature of noise in pre-training datasets and then effectively mitigate its impacts on downstream tasks. Specifically, through extensive experiments of fully-supervised and image-text contrastive pre-training on synthetic noisy ImageNet-1K, YFCC15M, and CC12M datasets, we demonstrate that, while slight noise in pre-training can benefit in-domain (ID) performance, where the training and testing data share a similar distribution, it always deteriorates out-of-domain (OOD) performance, where training and testing distributions are significantly different. These observations are agnostic to scales of pre-training datasets, pre-training noise types, model architectures, pre-training objectives, downstream tuning methods, and downstream applications. We empirically ascertain that the reason behind this is that the pre-training noise shapes the feature space differently. We then propose a tuning method (NMTune) to affine the feature space to mitigate the malignant effect of noise and improve generalization, which is applicable in both parameter-efficient and black-box tuning manners. We additionally conduct extensive experiments on popular vision and language models, including APIs, which are supervised and self-supervised pre-trained on realistic noisy data for evaluation. Our analysis and results demonstrate the importance of this novel and fundamental research direction, which we term as Noisy Model Learning.

Comments:	18 pages, 10 figures, 6 tables, preprint. arXiv admin note: substantial text overlap with arXiv:2309.17002
Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2403.06869 [cs.LG]
	(or arXiv:2403.06869v3 [cs.LG] for this version)
	http://doi.org/10.48550/arXiv.2403.06869

Submission history

From: Hao Chen [view email]
[v1] Mon, 11 Mar 2024 16:22:41 UTC (11,311 KB)
[v2] Fri, 14 Mar 2025 22:46:43 UTC (11,453 KB)
[v3] Mon, 5 May 2025 03:07:00 UTC (11,453 KB)

Computer Science > Machine Learning

Title:Impact of Noisy Supervision in Foundation Model Learning

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Impact of Noisy Supervision in Foundation Model Learning

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators