融合零样本语义先验与双几何约束的煤矿井下动态SLAM方法

Dynamic SLAM Method for Underground Coal Mines Integrating Zero-Shot Semantic Priors and Dual Geometric Constraints

  • 摘要: 针对煤矿井下场景语义标注数据匮乏、深度信息缺失以及人员、设备等动态目标运动模式复杂多样,导致视觉SLAM(Visual Simultaneous Localization and Mapping)在动态环境下特征误匹配增多、位姿估计稳定性下降的问题,提出一种融合零样本语义先验与双几何约束的煤矿井下动态SLAM方法。首先,构建基于开放词汇检测与光流提示的零样本语义分割模块,生成潜在动态目标掩码,并通过形态学处理优化边界,在无需井下专用标注数据的条件下为动态特征检测提供语义先验;其次,针对井下高反光、粉尘遮挡等因素造成的深度空洞问题,采用边缘优先填充的自适应深度图像修复策略补全深度缺失区域,增强深度信息完整性与几何约束可靠性。在此基础上,融合对极几何与深度重投影一致性约束,从二维匹配一致性和三维深度一致性两个方面联合判别动态特征,实现不同运动模式下动态特征的鲁棒识别与剔除,仅利用静态特征进行位姿估计。最后,在公开TUM RGB-D动态序列和自主采集的煤矿井下典型场景中开展实验。结果表明,所提方法能够有效剔除动态特征,减少特征误匹配,提升动态环境下的位姿估计精度和轨迹稳定性;在TUM RGB-D动态序列中,相较于DynaSLAM,所提方法的ATE RMSE和S.D.平均分别降低约23.92%和30.03%,且相较多种动态SLAM方法整体保持相近或更低的误差;在井下实景实验中,该方法轨迹波动和局部漂移较小,表现出较好的场景适用性与定位稳定性。

     

    Abstract: To address the problems of scarce semantic annotations, incomplete depth information, and complex motion patterns of dynamic objects such as personnel and equipment in underground coal mine scenes, which increase feature mismatches and reduce pose estimation stability in visual SLAM, this paper proposes a dynamic SLAM method integrating zero-shot semantic priors and dual geometric constraints. First, a zero-shot semantic segmentation module based on open-vocabulary detection and optical-flow prompts is constructed to generate potential dynamic object masks, whose boundaries are further refined by morphological processing, providing semantic priors for dynamic feature detection without mine-specific annotated data. Second, an edge-prior adaptive depth inpainting strategy is introduced to repair depth holes caused by high reflectance and dust occlusion, thereby improving depth completeness and the reliability of geometric constraints. On this basis, epipolar geometry and depth reprojection consistency are jointly employed to identify dynamic features from both two-dimensional matching consistency and three-dimensional depth consistency, enabling robust detection and removal of dynamic features under different motion patterns. Pose estimation is then performed using only static features. Experiments are conducted on public TUM RGB-D dynamic sequences and self-collected typical underground coal mine scenes. The results show that the proposed method effectively removes dynamic features, reduces feature mismatches, and improves pose estimation accuracy and trajectory stability in dynamic environments. On the TUM RGB-D dynamic sequences, compared with DynaSLAM, the proposed method reduces ATE RMSE and S.D. by approximately 23.92% and 30.03% on average, respectively, while maintaining comparable or lower errors than several dynamic SLAM methods. In real underground mine experiments, the method achieves smaller trajectory fluctuations and local drift, demonstrating good scene adaptability and localization stability.

     

/

返回文章
返回