Banner Banner

Deep Learning for Earth Observation: On Feature Reliance, Representation Learning and Label Noise

Tom Burgert

June 30, 2026

Technological advances in Earth observation (EO) have enabled frequent and high resolution monitoring of Earth at global scale, providing an unprecedented amount of remote sensing (RS) data. In recent years, deep learning methods originally devel oped in computer vision (CV) have been widely adopted to RS. While this adoption has enabled substantial progress, it relies on assumptions inherited from CV that are not always explicitly examined in the context of RS. Understanding where these assumptions hold and where they break down is essential for developing robust and interpretable learning methods for RS. To address this challenge in a structured manner, the thesis is organized around three complementary pillars that examine how domain-specific differences manifest across data characteristics (i.e., class fea tures), data properties (i.e., metadata), and task formulations. The first pillar of this thesis examines how data characteristics influence feature reliance in deep neural networks. Revisiting the long-standing hypothesis that ImageNet-trained convolu tional neural networks (CNNs) are inherently texture-biased, the thesis introduces a domain-agnostic evaluation framework based on controlled feature suppression. Empirical results from human and model experiments demonstrate that CNNs pre dominantly rely on local shape features instead of texture features. At the same time, systematic differences across domains are observed, with models trained on CVimages, medical images, and RS images exhibiting distinct reliance patterns on shape, color, and texture, respectively. These findings indicate that RS constitutes a distinct visual domain with its own feature hierarchies and inductive biases. The second pillar of this thesis addresses data properties specific to RS, focusing on the availability of auxiliary geographical metadata for contrastive self-supervised learning (SSL). To effectively exploit this information, the thesis introduces GeoRank, a regularization method for contrastive SSL that embeds geographic relationships directly into the representation space using spherical distance-based ranking. In addition to the methodological contribution, a systematic analysis of contrastive iii SSL design choices reveals that several assumptions commonly transferred from CV, including data augmentation strategies, dataset scaling, and input image size, do not consistently hold for multispectral RS images. The results highlight the need to align SSL objectives and induced invariances with the structure of RS data. The third pillar of this thesis focuses on task formulations, addressing multi-label classifica tion (MLC) under noisy supervision in RS. Multi-label datasets in RS often exhibit label noise due to cost-effective annotation procedures, such as thematic products or crowdsourcing. However, unlike single-label classification (SLC), noise in MLC can manifest in structurally different forms. To systematically analyze these effects, the thesis formalizes a taxonomy that distinguishes between additive, subtractive, and mixed multi-label noise and demonstrates that these noise types affect learning behavior in systematically different ways. Building on this analysis, NAR is intro duced as a noise-adaptive regularization method that adjusts supervision strength according to noise type and label entry reliability, substantially improving robustness across diverse noise conditions. Further, the thesis shows that the data augmenta tion technique CutMix can introduce multi-label noise in MLC and addresses this issue through a label propagation (LP) strategy. Taken together, the contributions of this thesis demonstrate that effective learning in RS requires explicitly account ing for domain-specific data characteristics, the effective exploitation of metadata, and domain-specific task formulations. By revisiting assumptions inherited from CV and introducing domain-specific learning strategies, the thesis advances the methodological foundations of deep learning for RS.

BIFOLD AUTHORS