文本检测方法、装置、电子设备、存储介质及产品

By employing a text detection method based on a data retrieval model, the semantic similarity and label similarity of the text to be detected are calculated, and noisy text is filtered out. This solves the problem of filtering low-quality data in the model training dataset and improves the accuracy of the model.

CN117149953BActive Publication Date: 2026-07-17ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
Filing Date
2023-08-29
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently filter out large amounts of low-quality data from model training datasets, leading to a decline in model accuracy.

Method used

By retrieving the text to be detected based on a data retrieval model, calculating the semantic similarity and/or text tag similarity of the text, filtering out noisy text, and obtaining the comparison text for text detection.

Benefits of technology

This enables efficient detection of low-quality data, improves the overall quality of the dataset, and thus enhances the accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117149953B_ABST
    Figure CN117149953B_ABST
Patent Text Reader

Abstract

本说明书实施例提供了一种文本检测方法、装置、电子设备、计算机可读存储介质及计算机程序产品,该方法包括:基于数据检索模型对待检测文本进行检索,得到待检测文本的至少一个近邻文本;确定待检测文本与至少一个近邻文本的文本语义相似度和 / 或文本标签相似度;基于文本语义相似度和 / 或文本标签相似度,对至少一个近邻文本中的噪声文本进行过滤,得到对比文本;基于对比文本对待检测文本进行文本检测。
Need to check novelty before this filing date? Find Prior Art