基于音形融合预训练的简帛文献补全方法
CSTR:
作者:
作者单位:

(1.武汉大学计算机学院,湖北武汉 430072;2.武汉大学文化遗产智能计算实验室,湖北武汉 430072;3.武汉大学信息管理学院,湖北武汉 430072;4.武汉大学历史学院,湖北武汉 430072)

作者简介:

刘嘉成(2001—),男,武汉大学计算机学院在读硕士研究生,研究方向为自然语言处理,E-mail: liu-jia-cheng@whu.edu.cn 通信作者:钱铁云(1970—),女,博士,武汉大学计算机学院教授,博士生导师,主要研究方向为数据库、自然语言处理与社会媒体处理,E-mail: qty@whu.edu.cn

通讯作者:

中图分类号:

基金项目:


Pretrained ancient text completion method based on phonetic and character shape fusion
Author:
Affiliation:

(1. School of Computer Science, Wuhan University, Wuhan 430072, China;2. Intellectual Computing Laboratory for Cultural Heritage, Wuhan University, Wuhan 430072, China;3. School of Information Management, Wuhan University, Wuhan 430072, China;4. School of History, Wuhan University, Wuhan 430072, China)

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    补全古文数字文本中的残缺内容,是提升简帛文献可读性与研究价值的重要技术路径。该研究提出了一种基于音形融合的预训练模型,通过拆解汉字的字形和发音,生成字符子结构,并利用交叉注意力机制整合这些信息,从而增强模型对古文语义的理解,提高简帛文献补全的准确性。具体方法包括利用五笔和注音对汉字进行建模,将每个汉字转化为短序列编码,并通过子词分割算法构建词汇表。实验结果显示,该方法在传世文献和出土文献数据集上的补全准确率分别达到了70.02%和63.76%,为简帛文献的智能补全提供了有效的解决方案。

    Abstract:

    Completing missing portions of ancient Chinese texts is a critical task in textual research and cultural heritage preservation. This study proposes a pretraining model based on the fusion of phonetic and character shape features to improve the accuracy of ancient text completion. The proposed method decomposes the shape and pronunciation of Chinese characters to generate sub-character structures and integrates this information through a cross-attention mechanism, thereby enhancing the model’s semantic understanding of ancient texts. The methodology involves modelling Chinese characters using Wubi input codes and phonetic systems, encoding each character as a short sequence, and constructing a vocabulary through a sub-word segmentation algorithm. Experimental results show that the proposed method achieves a completion accuracy of 70.02% on the transmitted literature dataset and 63.76% on the unearthed literature dataset, demonstrating its effectiveness for the intelligent completion of ancient texts.

    参考文献
    相似文献
    引证文献
引用本文

刘嘉成,李永奇,王晓光,李静,彭智勇,钱铁云.基于音形融合预训练的简帛文献补全方法[J].文物保护与考古科学,2026,(1):104-111.

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2024-12-07
  • 最后修改日期:2025-04-30
  • 录用日期:
  • 在线发布日期: 2026-03-12
  • 出版日期:
文章二维码
您是第位访问者
主办单位:上海博物馆 编辑出版:《文物保护与考古科学》编辑委员会
地址:上海市徐汇区龙吴路1118号,上海博物馆文物保护科技中心,《文物保护与考古科学》编辑部
电话:021-54362886 传真:021-54363740 E-mail:wwbhykgkx@163.com
文物保护与考古科学 ® 2026 版权所有
沪ICP备10003390号-3
沪公网安备 31010102005301号