Hybrid Transformer–Mamba Architecture for Topo-Bathymetric LiDAR Point Cloud Classification

Master Thesis at ifp - Mohammad Mahdi Baba

Mohammad Mahdi Baba

Duration: 6 months
Completion: May 2026
Supervisor: Dr.-Ing. Lida Asgharian Pournodrati
Examiner: Prof. Dr.-Ing. Uwe Sörgel

Introduction

Topo-bathymetric LiDAR provides detailed 3D point clouds across land–water boundaries and is highly valuable for shallow-water mapping, river monitoring, habitat analysis, and environmental applications. However, automatic classification of these data is challenging. The laser signal interacts with the water surface, water column, submerged vegetation, and bottom surface, which often creates noisy and ambiguous point distributions. In particular, separating aquatic plants from the seabed can be difficult because their geometry and radiometric responses may overlap in shallow or turbid water.

This thesis investigates a learning-based framework for classifying topo-bathymetric LiDAR point clouds into five semantic classes: water surface, seabed, aquatic plants, ground, and trees.

 

Method

The proposed workflow combines physically meaningful handcrafted features with efficient sequence modeling. First, the raw point cloud is organized through voxelization and hierarchical supervoxel segmentation. This provides stable local regions for feature extraction and reduces the influence of noise and irregular point density.

For each point, a compact 21-dimensional feature vector is constructed. It includes vertical structure cues, eigenvalue-based geometric descriptors, radiometric intensity statistics, and explicit dual-wavelength encoding for the green and near-infrared LiDAR channels. This is important because the green wavelength can penetrate the water column, while the near-infrared wavelength mainly captures water-surface and terrestrial returns.

To model local spatial context, the point cloud is divided into overlapping 3D windows. Each window is processed as a depth-ordered sequence using a Mamba-based encoder. Mamba was selected because it can model contextual dependencies with linear complexity, making it more efficient than attention-heavy Transformer models for dense point cloud windows. During inference, overlapping predictions are combined by majority voting to improve spatial consistency and reduce boundary errors.

Results

The method was evaluated on two large topo-bathymetric LiDAR datasets from Germany: Freiburg and Oppenheim. The Freiburg dataset represents a more structured case, while Oppenheim contains more complex aquatic conditions with stronger ambiguity between aquatic plants and seabed.

  • On the Freiburg dataset, the proposed method achieved an overall accuracy of 96.07% and a mean IoU of 92.30%.
  • On the Oppenheim dataset, it achieved an overall accuracy of 86.79% and a mean IoU of 77.21%.
  • The results show strong performance for water surface, ground, and trees, while the main remaining challenge is the separation of aquatic plants and seabed in complex underwater
Table 1: Quantitative results for various methods on the Freiburg dataset. The scores in bold indicate the best results among all methods.
Table 2: Quantitative results for various methods on the Oppenheim dataset. The scores in bold indicate the best results among all methods.

A comparison with a more complex Hybrid Transformer–Mamba model showed that the simpler Single-Mamba architecture provides a better trade-off between accuracy, robustness, and simplicity. The ablation study also confirmed that radiometric features and overlapping-window aggregation are important contributors to the final performance.

Conclusion

This thesis shows that topo-bathymetric LiDAR point cloud classification benefits from combining physically meaningful features, local overlapping context, and an efficient Mamba-based sequence model. The proposed framework produces accurate and spatially coherent classification results while avoiding the high computational cost of attention-based architectures.

The work also highlights the specific difficulty of shallow-water LiDAR classification: some errors are caused not only by model limitations, but also by genuine ambiguity in the data, especially where vegetation-like returns and seabed observations are physically mixed. Future work could improve the method further by adding depth-aware modeling, adaptive multi-scale windows, and waveform-based water-column information.

Ansprechpartner

Dieses Bild zeigt Uwe Sörgel

Uwe Sörgel

Prof. Dr.-Ing.

Institutsleiter, Fachstudienberater

Zum Seitenanfang