A text similarity calculation method and device based on a two-layer semantic model

The text similarity calculation method based on a two-layer semantic model solves the problem of semantic information loss in existing technologies, achieves accurate similarity calculation for both long and short texts, and is applicable to a variety of application scenarios.

CN115859993BActive Publication Date: 2026-07-24ZHEJIANG LAB +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG LAB
Filing Date
2022-11-08
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing text similarity calculation methods ignore or partially lose the semantic information of the text, resulting in reduced accuracy, especially when calculating the similarity between long and short texts.

Method used

A two-layer semantic model is adopted. First, the text sentence set is transformed into vectors through the first semantic model. Then, the sentence vector set is encoded through the second semantic model. Finally, the text similarity is calculated by combining the text length contrast.

Benefits of technology

It preserves the semantic information of the text to the greatest extent, expands the scope of application to similarity calculation of both long and short texts, improves the accuracy and breadth of the calculation, and is suitable for scenarios such as similar text matching and deduplication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115859993B_ABST
    Figure CN115859993B_ABST
Patent Text Reader

Abstract

The application discloses a text similarity calculation method and device based on a two-layer semantic model, counts the number of sentences of a first text and a second text, records the smaller number as a first text sentence set and the other as a second text sentence set, and calculates a text length contrast; the first text sentence set and the second text sentence set are respectively subjected to vector conversion through a first semantic model to obtain a first text sentence vector set and a second text sentence vector set; the distance similarity of each sentence vector is calculated to find the most similar sentence corresponding to each sentence of the first text sentence set in the second text sentence set; the most similar sentences are combined to obtain a third text sentence vector set; the first text sentence vector set and the third text sentence vector set are subjected to coding through a second semantic model to obtain a first text vector and a third text vector, and the similarity of the first text vector and the third text vector is calculated; the vector similarity is multiplied by the text length contrast to obtain the text similarity.
Need to check novelty before this filing date? Find Prior Art