Token hierarchical retrieval method and system based on vector similarity calculation

By employing a token-based hierarchical retrieval method to segment, hierarchically divide, and vectorize sentences and enterprise-level data, and combining this with cosine similarity calculation, the problem of insufficient distinguishability caused by high-density data in enterprise-level retrieval is solved, thereby improving retrieval accuracy and efficiency.

CN120821829APending Publication Date: 2025-10-21BEIJING NEUSOFT HUIJU INFORMATION TECH HLDG CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511318048.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

In enterprise-level vector retrieval scenarios, the high density of similar data leads to insufficient discriminative power. Existing methods are not sensitive to subtle semantic differences. In particular, the calculation of Euclidean distance and cosine similarity in high-dimensional space suffers from the 'curse of dimensionality', resulting in denser search results and making it difficult to accurately distinguish semantically similar documents.

Method used

The token-based hierarchical retrieval method is adopted, which involves word segmentation, token layering, vectorization, and cosine similarity calculation of the query statement and enterprise-level data. The hierarchical processing is combined with the importance of the enterprise-level data to improve semantic distinguishability.

Benefits of technology

By using token-based hierarchical processing, the accuracy of retrieval has been significantly improved, ensuring that semantically similar documents can be more clearly distinguished in the search results. The dynamic adjustment mechanism ensures that the hierarchical strategy is synchronized with business development, thereby improving the accuracy and efficiency of retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120821829A_ABST
    Figure CN120821829A_ABST
Patent Text Reader

Abstract

The invention relates to the cross technical field of artificial intelligence and information retrieval, and discloses a Token hierarchical retrieval method and system based on vector similarity calculation, and the method comprises the steps: S1, carrying out the word segmentation of a to-be-retrieved statement and enterprise-level data, and obtaining retrieval word segmentation data and enterprise word segmentation data; step S2, performing Token layering on the basis of the retrieval word segmentation data and the enterprise word segmentation data to obtain Token layering data; step S3, carrying out vectorization calculation on the Token hierarchical data to obtain a Token hierarchical vector; and S4, calculating cosine similarity based on the Token hierarchical vector to obtain enterprise-level data related to the statement to be retrieved. According to the method, the business correlation is adopted as a core layering basis, and the Token and the enterprise business process are deeply bound, so that the business key terms obtain higher weights.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intersectional technology of artificial intelligence and information retrieval, and in particular to a token hierarchical retrieval method and system based on vector similarity calculation. Background Art

[0002] In the fields of information retrieval and machine learning, vector similarity calculation is a core foundational technology, and its algorithmic evolution has formed a complete system. Current mainstream methods are based on two paradigms: 1) Exact distance metrics (such as Euclidean distance and cosine similarity), which determine vector relevance through geometric spatial relationships; and 2) Approximate nearest neighbor search (such as Locality Sensitive Hashing (LSH) and Product Quantization (PQ), which trades precision for efficiency to achieve large-scale retrieval. These methods all require normalizing the input vector (constraining the range to the [0, 1] range). Their dimensionality reduction visualization reveals a typical characteristic: on a two-dimensional plane, the data points are distributed in a ring-shaped region centered at the origin (a hollow circle with a radius of approximately 1.0). Existing optimization research is largely based on this geometric characteristic, including methods such as density redistribution of the ring region, dynamic radius adjustment, and nonlinear deformation based on manifold learning.

[0003] In enterprise-level vector retrieval scenarios, insufficient discrimination due to densely similar data is a typical challenge. Normalized vectors exhibit a hyperspherical distribution in high-dimensional space. When enterprise-level data has strong domain characteristics, semantically similar documents tend to cluster in a narrow area, resulting in minimal differences in cosine similarity and Euclidean distance values. Tests have shown that in a knowledge base with tens of millions of entries, the similarity difference between the top 1,000 results is often less than 0.08.

[0004] Euclidean distance and cosine similarity are insensitive to subtle semantic differences. Especially when the vector dimension exceeds 512, the distance calculation suffers from the "curse of dimensionality", which exacerbates the density of the results.

[0005] There are many existing solutions: first, using a domain-adapted embedding model to enhance semantic differentiation; second, using a hybrid search strategy combined with a keyword filtering pre-screening set to perform vector-based precision ranking of search results; third, dynamic threshold adjustment technology automatically expands the TOP-K range based on the query vector density, combined with secondary clustering screening; fourth, distance metric optimization uses manifold learning (such as t-SNE) to perform local space stretching and amplify key feature differences. Summary of the Invention

[0006] In order to solve the above technical problems, the present invention proposes a token hierarchical retrieval method and system based on vector similarity calculation, which specifically includes:

[0007] Step S1: Segment the search statement and enterprise-level data to obtain search segmentation data and enterprise segmentation data;

[0008] Step S2: Token layering is performed based on the search word segmentation data and the enterprise word segmentation data to obtain token layering data;

[0009] Step S3: performing vectorized calculation on the Token layered data to obtain a Token layered vector;

[0010] Step S4: Calculate the cosine similarity based on the Token hierarchical vectors to obtain enterprise-level data related to the sentence to be retrieved.

[0011] Optionally, in step S2, the search word segmentation data and the enterprise word segmentation data are tokenized according to the importance of the enterprise-level data matching query to obtain token-layered data.

[0012] Optionally, in step S4, the cosine similarity is calculated as follows:

[0013] ;

[0014] in, is the angle between vector A and the horizontal coordinate axis, is the angle between vector B and the horizontal coordinate axis, is the length of the adjacent side of vector A, is the length of the adjacent side of vector B, is the length of the opposite side of vector A, is the length of the opposite side of vector B, is the length of the hypotenuse of vector A, is the length of the hypotenuse of vector B.

[0015] The present invention also discloses a Token hierarchical retrieval system based on vector similarity calculation, the system comprising:

[0016] The word segmentation module is used to segment the search statement and enterprise-level data to obtain search word segmentation data and enterprise word segmentation data;

[0017] A Token stratification module, configured to perform Token stratification based on the search word segmentation data and the enterprise word segmentation data to obtain Token stratification data;

[0018] A vectorization module, configured to perform vectorization calculation on the Token layered data to obtain a Token layered vector;

[0019] The similarity calculation module is used to calculate the cosine similarity based on the Token hierarchical vector to obtain enterprise-level data related to the sentence to be retrieved.

[0020] Optionally, the search word segmentation data and the enterprise word segmentation data are tokenized according to the importance of the enterprise-level data matching query to obtain token-layered data.

[0021] Optionally, the cosine similarity is calculated as:

[0022] ;

[0023] in, is the angle between vector A and the horizontal coordinate axis, is the angle between vector B and the horizontal coordinate axis, is the length of the adjacent side of vector A, is the length of the adjacent side of vector B, is the length of the opposite side of vector A, is the length of the opposite side of vector B, is the length of the hypotenuse of vector A, is the length of the hypotenuse of vector B.

[0024] Compared with the prior art, the present invention has the following beneficial effects:

[0025] During the AI ​​project's development, token stratification served as a core preprocessing step, initially employing traditional part-of-speech classification (eight major word classes, including nouns, verbs, and adjectives). However, in actual applications, it was found that nouns accounted for a high proportion of enterprise-level data, leading to a "noun clustering effect" in the stratification results, which affected subsequent feature extraction.

[0026] Currently, we use business relevance as the core tiering basis, deeply binding tokens to enterprise business processes and assigning higher weights to key business terms. This adjustment has already shown positive results, with search accuracy steadily improving. We will continue to optimize tiering effectiveness through a dynamic adjustment mechanism, including establishing a regular review process and an automated assessment system, to ensure that the tiering strategy evolves in tandem with business development. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0028] Figure 1 This is a method step diagram of a token hierarchical retrieval method based on vector similarity calculation provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0029] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0030] Explanation of proper nouns:

[0031] Word segmentation: dividing text information into basic units (words, characters);

[0032] Vectorization: Mapping the basic units after word segmentation to the vector library through the Embedding model is a feature extraction step;

[0033] Token: The basic processing unit in natural language processing (NLP) and large language models (LLM). It refers to the smallest meaningful segment of text after segmentation. The segmentation process is called word segmentation.

[0034] Token layering: This is a technical step that performs structured processing on basic word segmentation units (tokens). It aims to inject multi-dimensional features into subsequent vectorization and enhance the depth and precision of semantic representation.

[0035] The traditional calculation method is as follows:

[0036] (1) Participle:

[0037] Sentence A: Supplier / certificate / what / when / expires;

[0038] Sentence B: Supplier / Master Data / Includes / Qualification / Certificate / Expiration / Time.

[0039] (2) Vectorization: as shown in Table 1.

[0040] Table 1

[0041]

[0042] (3) Calculate using the cosine similarity formula:

[0043] .

[0044] Example 1

[0045] like Figure 1 As shown in the figure, a token hierarchical retrieval method based on vector similarity calculation mainly includes:

[0046] Step S1: Segment the search statement and enterprise-level data to obtain search segmentation data and enterprise segmentation data.

[0047] Step S2: Token stratification is performed based on the search word segmentation data and the enterprise word segmentation data to obtain token stratification data. Specifically, the search word segmentation data and the enterprise word segmentation data are token stratified according to the importance of the enterprise-level data matching query to obtain token stratification data.

[0048] Step S3: perform vectorized calculation on the Token layered data to obtain the Token layered vector.

[0049] Step S4: Calculate the cosine similarity based on the Token hierarchical vector to obtain enterprise-level data related to or similar to the query statement. Specifically, the cosine similarity calculation method is:

[0050] ;

[0051] in, is the angle between vector A and the horizontal coordinate axis, is the angle between vector B and the horizontal coordinate axis, is the length of the adjacent side of vector A, is the length of the adjacent side of vector B, is the length of the opposite side of vector A, is the length of the opposite side of vector B, is the length of the hypotenuse of vector A, is the length of the hypotenuse of vector B.

[0052] Example 2:

[0053] Token hierarchical retrieval method based on vector similarity calculation, such as Figure 1 As shown, specifically including:

[0054] Step S1: Segment the search statement and enterprise-level data to obtain search segmentation data and enterprise segmentation data.

[0055] Sentence A: Supplier / certificate / what / when / expires;

[0056] Sentence B: Supplier / Master Data / Includes / Qualification / Certificate / Expiration / Time.

[0057] In this embodiment, the company extracts enterprise knowledge base data through the baseline system and fine-tunes the Qwen3-1.7B model based on this data. In the inference stage, the fine-tuned model input (prompt words + user questions) is automatically segmented by the large model.

[0058] Step S2: Token stratification is performed based on the search word segmentation data and the enterprise word segmentation data to obtain Token stratification data.

[0059] Tokens are stratified according to the importance of enterprise-level data matching queries. The results after stratification are shown in Table 2:

[0060] Table 2

[0061]

[0062] Step S3: perform vectorized calculation on the Token layered data to obtain a Token layered vector.

[0063] Vector encoding upgrade: The original 2048-dimensional binary vector (such as (1,0,1,1,...)) is transformed through token layering, replacing the single-bit (0 / 1) unit with the two-bit combination unit [0,0], [0,1], [1,0], [1,1], generating a new vector example ([1,0], [0,1],...).

[0064] The token layered vector results are shown in Table 3:

[0065] Table 3

[0066]

[0067] Step S4: Calculate the cosine similarity based on the Token hierarchical vectors to obtain enterprise-level data related to or similar to the sentence to be retrieved.

[0068] The cosine similarity in two-dimensional space is extended to the cosine similarity in n-dimensional space. The calculation method of cosine similarity is:

[0069] .

[0070] Judging from the two results, the similarity has improved. When combined with the existing similarity matching framework, the query results will have a more obvious improvement.

[0071] Example 3

[0072] A token hierarchical retrieval system based on vector similarity calculation, mainly including:

[0073] The word segmentation module is used to segment the search statement and enterprise-level data to obtain search word segmentation data and enterprise word segmentation data.

[0074] The Token stratification module is used to perform Token stratification based on the search word segmentation data and the enterprise word segmentation data to obtain Token stratification data. Specifically, the Token stratification is performed on the search word segmentation data and the enterprise word segmentation data according to the importance of the enterprise-level data matching query to obtain Token stratification data.

[0075] The vectorization module is used to perform vectorized calculations on Token layered data to obtain Token layered vectors.

[0076] The similarity calculation module is used to calculate the cosine similarity based on the Token layer vector to obtain enterprise-level data related to or similar to the query statement. Specifically, the cosine similarity calculation method is:

[0077] ;

[0078] in, is the angle between vector A and the horizontal coordinate axis, is the angle between vector B and the horizontal coordinate axis, is the length of the adjacent side of vector A, is the length of the adjacent side of vector B, is the length of the opposite side of vector A, is the length of the opposite side of vector B, is the length of the hypotenuse of vector A, is the length of the hypotenuse of vector B.

[0079] Example 4:

[0080] A token hierarchical retrieval system based on vector similarity calculation, which includes:

[0081] The word segmentation module is used to segment the search statement and enterprise-level data to obtain search word segmentation data and enterprise word segmentation data.

[0082] Sentence A: Supplier / certificate / what / when / expires;

[0083] Sentence B: Supplier / Master Data / Includes / Qualification / Certificate / Expiration / Time.

[0084] The Token stratification module is used to perform Token stratification based on the search word segmentation data and the enterprise word segmentation data to obtain Token stratification data.

[0085] Tokens are stratified according to the importance of enterprise-level data matching queries. The results after stratification are shown in Table 4:

[0086] Table 4

[0087]

[0088] The vectorization module is used to perform vectorized calculation on the Token layered data to obtain a Token layered vector.

[0089] Vector encoding upgrade: The original 2048-dimensional binary vector (such as (1,0,1,1,...)) is transformed through token layering, replacing the single-bit (0 / 1) unit with the two-bit combination unit [0,0], [0,1], [1,0], [1,1], generating a new vector example ([1,0], [0,1],...).

[0090] The token layered vector results are shown in Table 5:

[0091] Table 5

[0092]

[0093] The similarity calculation module is used to calculate the cosine similarity based on the Token hierarchical vector to obtain enterprise-level data related to or similar to the sentence to be retrieved.

[0094] The calculation method of cosine similarity is:

[0095] .

[0096] Judging from the two results, the similarity has improved. When combined with the existing similarity matching framework, the query results will have a more obvious improvement.

[0097] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.

Claims

1. A token hierarchical retrieval method based on vector similarity calculation, characterized in that: The method comprises: Step S1: Segment the search statement and enterprise-level data to obtain search segmentation data and enterprise segmentation data; Step S2: Token layering is performed based on the search word segmentation data and the enterprise word segmentation data to obtain token layering data; Step S3: performing vectorized calculation on the Token layered data to obtain a Token layered vector; Step S4: Calculate the cosine similarity based on the Token hierarchical vectors to obtain enterprise-level data related to the sentence to be retrieved.

2. The Token hierarchical retrieval method based on vector similarity calculation according to claim 1 is characterized in that: In the step S2, the search word segmentation data and the enterprise word segmentation data are tokenized according to the importance of the enterprise-level data matching query to obtain token-layered data.

3. The Token hierarchical retrieval method based on vector similarity calculation according to claim 2 is characterized in that: In step S4, the cosine similarity is calculated as follows: ; in, is the angle between vector A and the horizontal coordinate axis, is the angle between vector B and the horizontal coordinate axis, is the length of the adjacent side of vector A, is the length of the adjacent side of vector B, is the length of the opposite side of vector A, is the length of the opposite side of vector B, is the length of the hypotenuse of vector A, is the length of the hypotenuse of vector B.

4. A token hierarchical retrieval system based on vector similarity calculation, the system being used to implement the retrieval method according to any one of claims 1 to 3, characterized in that: The system includes: The word segmentation module is used to segment the search statement and enterprise-level data to obtain search word segmentation data and enterprise word segmentation data; A Token stratification module, configured to perform Token stratification based on the search word segmentation data and the enterprise word segmentation data to obtain Token stratification data; A vectorization module, configured to perform vectorization calculation on the Token layered data to obtain a Token layered vector; The similarity calculation module is used to calculate the cosine similarity based on the Token hierarchical vector to obtain enterprise-level data related to the sentence to be retrieved.

5. The Token hierarchical retrieval system based on vector similarity calculation according to claim 4 is characterized in that: The search word segmentation data and the enterprise word segmentation data are tokenized according to the importance of the enterprise-level data matching query to obtain token-layered data.

6. The Token hierarchical retrieval system based on vector similarity calculation according to claim 5 is characterized in that: The calculation method of cosine similarity is: ; in, is the angle between vector A and the horizontal coordinate axis, is the angle between vector B and the horizontal coordinate axis, is the length of the adjacent side of vector A, is the length of the adjacent side of vector B, is the length of the opposite side of vector A, is the length of the opposite side of vector B, is the length of the hypotenuse of vector A, is the length of the hypotenuse of vector B.

Citation Information

Patent Citations

  • Text similarity matching and calculating method, system and device

    CN112364124A

  • Enterprise document library construction and retrieval method and system

    CN117421333A

  • Intelligent search engine system oriented to enterprise intranet

    CN117891804A

  • Enterprise information generation and retrieval method based on AI and knowledge graph

    CN120578742A