Knowledge base system for matching vectorized knowledge representation and large model

Through a knowledge base system that matches large models with vectorized knowledge representation, the problem of low computation and matching accuracy of knowledge base systems in vector space in the prior art is solved, and efficient and personalized knowledge retrieval services are realized.

CN120470033AActive Publication Date: 2025-08-12GUANGZHOU JIAYIN COMMUNICATION CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510645253.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-08-12
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

The existing knowledge base system is difficult to effectively calculate and match in vector space, and the lack of reasonable weighting of matching scores based on the domain knowledge graph, resulting in low matching retrieval accuracy between user queries and knowledge entries.

Method used

A knowledge base system that uses vectorized knowledge to represent matching with large models, including structured analysis module, knowledge entry vectorization module, vector multi-dimensional matching module and comprehensive weight matching module. Through structured analysis, vectorization and multi-dimensional matching scores, the matching retrieval accuracy of user queries and knowledge entries is improved.

Benefits of technology

It realizes effective calculation and matching between user query and knowledge entries, improves matching search accuracy, generates personalized search reports, meets users' query needs in specific fields, and improves the efficiency and effectiveness of knowledge acquisition and utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470033A_ABST
    Figure CN120470033A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of knowledge matching retrieval, and particularly discloses a vectorized knowledge representation and large-scale model matching knowledge base system, which comprises a structured analysis module used for obtaining all structured comprehensive character knowledge entries in a knowledge base; the knowledge entry vectorization module is used for converting each structured comprehensive character knowledge entry into a high-dimensional vector based on a vectorization model; the vector multi-dimensional matching module is used for generating a user query vector and calculating a multi-dimensional matching score of the user query vector and the high-dimensional vector of each comprehensive character knowledge item in the knowledge base; the comprehensive weight matching module is used for performing weight endowing on the multi-dimensional matching score based on semantics and position constraint degrees of all synonym entity combinations in the knowledge graph of the domain to which the user queries belong, obtaining a matching result based on the multi-dimensional matching score and the multi-dimensional endowing weight, and generating a personalized retrieval report based on the matching result; the method and the device are used for improving knowledge matching and retrieval precision and providing high-quality retrieval results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge matching and retrieval, and in particular to a knowledge base system for matching vectorized knowledge representation with large models. Background Art

[0002] In the era of information explosion, knowledge base systems are crucial for the storage, management, and retrieval of knowledge. With the rapid growth of data volumes and the increasing complexity of user needs, traditional knowledge base systems face numerous challenges. Accurately and efficiently retrieving information that matches user needs from vast amounts of knowledge has become a pressing issue. Vectorized knowledge representation and large-scale model matching technologies offer new opportunities for optimizing knowledge base systems. Converting knowledge items into vector form allows for better utilization of the mathematical properties of vector spaces for similarity calculation and matching, improving retrieval accuracy and efficiency. Furthermore, combined with large-scale models, these systems can fully leverage their advantages in semantic understanding and knowledge reasoning, deeply exploring relationships between knowledge and providing users with more intelligent and precise knowledge services. This combined knowledge base system holds broad application prospects in numerous fields, such as intelligent question answering, information retrieval, and intelligent decision-making, and is expected to propel knowledge management and utilization to new heights.

[0003] However, existing knowledge base systems have several flaws. Knowledge is difficult to effectively compute and match in vector space. The lack of an ability to appropriately weight matching scores based on domain knowledge graphs affects the accuracy of matching retrieval between user queries and knowledge items, making it difficult for users to quickly obtain accurate and relevant knowledge information.

[0004] Therefore, the present invention proposes a knowledge base system that matches vectorized knowledge representation with large models. Summary of the Invention

[0005] The present invention provides a knowledge base system that matches vectorized knowledge representation with large-scale models, which is used to realize the effective operation and matching of user queries and knowledge items in the vector space, introduce the knowledge graph of the field to which the user query belongs, and realize the accurate and effective weighting of the multi-dimensional matching scores of user queries and knowledge items in the vector space, thereby improving the matching retrieval accuracy between user queries and knowledge items and providing high-quality retrieval results.

[0006] The present invention provides a knowledge base system for matching vectorized knowledge representation with large models, comprising:

[0007] A structured parsing module is used to perform structured parsing on all unstructured comprehensive character knowledge items in the knowledge base to obtain all structured comprehensive character knowledge items;

[0008] A knowledge item vectorization module is used to convert each structured comprehensive character knowledge item into a high-dimensional vector based on a vectorization model, and obtain a comprehensive character knowledge item high-dimensional vector of each structured comprehensive character knowledge item;

[0009] Vector multi-dimensional matching module, used to generate user query vectors and calculate the multi-dimensional matching scores between the user query vectors and the high-dimensional vectors of each comprehensive character knowledge item in the knowledge base;

[0010] The comprehensive weight matching module is used to assign weights to multidimensional matching scores based on the semantics and position constraints of all synonymous entity combinations in the knowledge graph of the user's query field, obtain multidimensional weights of the multidimensional matching scores, obtain matching results based on the multidimensional matching scores and the multidimensional weights, and generate personalized retrieval reports based on the matching results.

[0011] Preferably, the knowledge item vectorization module includes:

[0012] The knowledge item encoding submodule is used to perform absolute position encoding and relative position association encoding on all sub-information bodies in each structured comprehensive character knowledge item, and generate the position and semantic perception vector of each sub-information body;

[0013] The feature alignment submodule is used to dimensionally align the positions and semantic perception vectors of all character class sub-information bodies in each structured comprehensive character class knowledge entry and map them to the same feature space to obtain a high-dimensional vector of the comprehensive character class knowledge entry for each structured comprehensive character class knowledge entry.

[0014] Preferably, the knowledge item encoding submodule includes:

[0015] A basic coding unit is used to divide all character class sub-information bodies in each structured comprehensive character class knowledge entry, encode each sub-information body based on its semantic content, and obtain a basic code for each sub-information body;

[0016] An absolute position encoding unit, for generating an absolute position encoding of each sub-information body based on the different character class positions and the same character class positions of each sub-information body;

[0017] A relative position encoding unit, configured to generate a relative position association code for each sub-information body based on the relative association positions of the same character class and the relative association positions of the different character classes of each sub-information body;

[0018] The position and semantic perception vectorization unit is used to generate the position and semantic perception vector of each sub-information body based on the basic coding, absolute position coding, relative position coding and vectorization model of each sub-information body.

[0019] Preferably, the absolute position encoding unit includes:

[0020] The variant character class positioning subunit is used to determine the sorting position of each sub-information body in the corresponding structured comprehensive character class knowledge entry, and generate the variant character class position of each sub-information body based on the sorting position of each sub-information body in the corresponding structured comprehensive character class knowledge entry;

[0021] The information body classification sub-unit is used to classify the sub-information bodies of the same character class in all the sub-information bodies of the character class in each structured comprehensive character class knowledge entry in an order-preserving manner to obtain a set of all sub-information bodies of the same character class in each structured comprehensive character class knowledge entry;

[0022] The same character class positioning subunit is used to determine the sorting position of each sub-information body in the set of sub-information bodies of the same character class to which it belongs, and generate the same character class position of each sub-information body based on the set of sub-information bodies of the same character class to which each sub-information body belongs and its sorting position in the set of sub-information bodies of the same character class to which it belongs;

[0023] The absolute position coding subunit is used to generate the absolute position coding of each sub-information body based on the different character class positions and the same character class positions of each sub-information body and the preset absolute position coding method.

[0024] Preferably, the relative position encoding unit includes:

[0025] The same-character-class relative association positioning subunit is used to treat the sorting position of all synonymous sub-information bodies of each sub-information body of each structured comprehensive character class knowledge item in the same-character-class sub-information body set as the same-character-class relative association position of each sub-information body;

[0026] The different character class relative association positioning subunit is used to use the sorting position of all synonymous sub-information bodies in the sub-information bodies corresponding to all character classes of each sub-information body of each structured comprehensive character class knowledge item as the different character class relative association position of each sub-information body;

[0027] The relative position coding subunit is used to generate a relative position association code for each sub-information body based on the relative associated positions of the same character class and the relative associated positions of different character classes of each sub-information body and a preset relative position association coding method.

[0028] Preferably, the vector multi-dimensional matching module includes:

[0029] The information body vectorization submodule is used to determine the semantic and position perception vectors of each sub-information body queried by the user based on the user query vector, and at the same time determine the semantic and position perception vectors of each sub-information body in each structured comprehensive character knowledge item based on the high-dimensional vector of each comprehensive character knowledge item in the knowledge base;

[0030] A semantic and position matching submodule is used to calculate the semantic and position matching score of each sub-information body queried by the user and each sub-information body in each structured comprehensive character knowledge entry based on the similarity of all elements in the same position in the semantic and position perception vector of each sub-information body queried by the user and the semantic and position perception vector of each sub-information body in each structured comprehensive character knowledge entry;

[0031] A matching score matrix submodule is used to construct a semantic and position matching matrix between the user query and each structured comprehensive character knowledge item based on the semantic and position matching scores of each sub-information body in the user query and each structured comprehensive character knowledge item;

[0032] The multidimensional matching submodule is used to construct a mapping curve based on all diagonal elements and all non-diagonal elements of the semantic and position matching matrix between the user query and each structured comprehensive character knowledge item, and calculate the multidimensional matching score between the user query vector and the high-dimensional vector of each comprehensive character knowledge item in the knowledge base.

[0033] Preferably, the multi-dimensional matching submodule includes:

[0034] A first mapping curve generating unit is configured to sort all elements greater than a first threshold in a semantic and position matching matrix between the user query and each structured comprehensive character knowledge item from largest to smallest to obtain a first element sequence, and generate a first mapping curve based on the row and column ordinal numbers of all elements in the first element sequence in the semantic and position matching matrix;

[0035] A second mapping curve generating unit is configured to sort all diagonal elements greater than a second threshold value in a semantic and position matching matrix between the user query and each structured comprehensive character knowledge item from largest to smallest to obtain a second element sequence, and generate a second mapping curve based on the row and column ordinal numbers of all elements in the second element sequence in the semantic and position matching matrix;

[0036] A third mapping curve generating unit is configured to sort all non-diagonal elements of the semantic and position matching matrix between the user query and each structured comprehensive character knowledge item that are greater than a third threshold value from largest to smallest to obtain a third element sequence, and generate a third mapping curve based on the row and column ordinal numbers of all elements in the third element sequence in the semantic and position matching matrix;

[0037] A multidimensional matching degree calculation unit is used to treat the matching degree between the first mapping curve and the first standard mapping curve, the matching degree between the second mapping curve and the second standard mapping curve, and the matching degree between the third mapping curve and the third standard mapping curve as the multidimensional matching scores of the user query vector and the high-dimensional vector of each comprehensive character knowledge item in the knowledge base.

[0038] Preferably, the comprehensive weight matching module includes:

[0039] The semantic and position constraint determination submodule is used to screen out all synonymous entity combinations in the knowledge graph of the domain to which the user queries belong, and determine the semantic and position constraint degrees of all synonymous entity combinations based on their ranking values in all related reference knowledge items;

[0040] A weighting submodule is used to weight the multidimensional matching scores based on the semantics and position constraints of all synonymous entity combinations to obtain multidimensional weighting of the multidimensional matching scores;

[0041] A calibration matching score generation submodule is used to obtain a calibration matching score between the user query and the corresponding structured comprehensive character knowledge item by weighted summation of the multi-dimensional matching scores based on the multi-dimensional weights assigned, and to use the calibration matching scores of all structured comprehensive character knowledge items in the knowledge base as matching results;

[0042] A first knowledge item screening submodule is configured to treat all structured comprehensive character knowledge items in the knowledge base whose calibration matching scores are greater than a first score threshold as all target knowledge items, and generate a semantic association graph between the user query and all structured comprehensive character knowledge items;

[0043] A second knowledge item screening submodule is configured to treat all structured comprehensive character knowledge items in the knowledge base whose calibration matching scores are greater than a second score threshold and not greater than the first score threshold as all expanded recommended knowledge items;

[0044] The personalized retrieval report generation submodule is used to generate a personalized retrieval report based on all target knowledge items and all expanded recommended knowledge items.

[0045] Preferably, the semantic and position constraint determination submodule includes:

[0046] The semantic and position constraint matrix generation submodule is used to filter out all synonymous entity combinations in the knowledge graph of the domain to which the user query belongs;

[0047] The semantic and position constraint calculation submodule is used to determine the semantic and position constraint of each synonymous entity combination based on the comprehensive similarity and proximity of the ranking values of each synonymous entity combination in all corresponding reference knowledge items.

[0048] Preferably, the weight assignment submodule includes:

[0049] a synonymous entity combination counting unit, configured to count the number of all synonymous entity combinations whose semantic and positional constraints are greater than a constraint threshold as a first number, and at the same time, count the number of all synonymous entity combinations whose semantic and positional constraints are not greater than the constraint threshold as a second number;

[0050] A multi-dimensional weight assigning unit, configured to obtain a multi-dimensional matching score based on the first quantity and the second quantity, comprising:

[0051] The weight of the matching degree between the first mapping curve and the first standard mapping curve is set to 0.5, and the product of the ratio of the first number to the total number of all synonymous entity combinations and 0.5 is used as the weight of the matching degree between the second mapping curve and the second standard mapping curve, and the product of the ratio of the second number to the total number of all synonymous entity combinations and 0.5 is used as the weight of the matching degree between the third mapping curve and the third standard mapping curve.

[0052] The beneficial effects of the present invention compared to the prior art are as follows: the structured parsing module performs structured parsing on the unstructured comprehensive character knowledge items in the knowledge base. This operation organizes the disorganized knowledge, making it more organized and easy to understand, greatly improving the processability and retrievability of knowledge, and facilitating the in-depth mining and utilization of knowledge by subsequent systems. The knowledge item vectorization module converts the structured knowledge items into high-dimensional vectors with the help of a vectorization model, representing knowledge in the form of vectors, which can more accurately characterize the intrinsic characteristics and semantic information of knowledge, facilitate efficient storage, calculation and analysis by computers, and provide a more accurate data foundation for knowledge matching. The vector multi-dimensional matching module generates a user query vector and calculates its multi-dimensional matching score with the high-dimensional vector of the knowledge item in the knowledge base, thereby evaluating the degree of matching between knowledge and user queries from multiple dimensions, making the matching process more comprehensive and detailed, improving the accuracy and rationality of the matching, and avoiding the one-sidedness of single-dimensional matching. The comprehensive weighted matching module assigns weights to multidimensional matching scores based on the semantics and positional constraints of synonymous entity combinations in the domain knowledge graph. This module fully considers the semantic relationships and structural characteristics of the knowledge domain to further optimize matching results, better meeting user query needs in specific domains and generating personalized search reports that better meet their actual needs. Through the collaborative work of various modules, the entire system achieves efficient operation of the entire process, from structured knowledge processing to vectorized representation, to precise matching and personalized report generation. This provides users with high-quality, personalized knowledge retrieval services, helps improve the efficiency and effectiveness of knowledge acquisition and utilization, and is of great value in the fields of knowledge management and application.

[0053] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in this application document.

[0054] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0056] Figure 1 A knowledge base system that matches vectorized knowledge representation with large models in an embodiment of the present invention;

[0057] Figure 2 4 is an architectural diagram of a knowledge item vectorization module according to an embodiment of the present invention;

[0058] Figure 3 4 is an architecture diagram of a vector multi-dimensional matching module in an embodiment of the present invention;

[0059] Figure 4 4 is an architectural diagram of the comprehensive weight matching module in an embodiment of the present invention. DETAILED DESCRIPTION

[0060] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0061] Example 1: The present invention provides a knowledge base system that matches vectorized knowledge representation with large models, referring to Figure 1 ,include:

[0062] A structured parsing module is used to perform structured parsing on all unstructured comprehensive character knowledge items in the knowledge base to obtain all structured comprehensive character knowledge items;

[0063] A knowledge item vectorization module is used to convert each structured comprehensive character knowledge item into a high-dimensional vector based on a vectorization model, and obtain a comprehensive character knowledge item high-dimensional vector of each structured comprehensive character knowledge item;

[0064] Vector multi-dimensional matching module, used to generate user query vectors and calculate the multi-dimensional matching scores between the user query vectors and the high-dimensional vectors of each comprehensive character knowledge item in the knowledge base;

[0065] The comprehensive weight matching module is used to assign weights to multidimensional matching scores based on the semantics and position constraints of all synonymous entity combinations in the knowledge graph of the user's query field, obtain multidimensional weights of the multidimensional matching scores, obtain matching results based on the multidimensional matching scores and the multidimensional weights, and generate personalized retrieval reports based on the matching results.

[0066] In this embodiment, the knowledge base is a collection of stored knowledge, covering various types of knowledge information. It is like a vast knowledge warehouse, providing data support for the entire knowledge matching and retrieval system. For example, a legal knowledge base may contain various legal provisions, legal case analysis, and other knowledge content. When a user conducts a legal query, the system will search this knowledge base for matching information.

[0067] In this embodiment, unstructured comprehensive character knowledge entries refer to a series of independent entries that lack a specific structure, are presented in the form of characters, and are relatively disorganized. These entries can provide detailed analysis and explanation of knowledge in a specific field, or they can provide a brief introduction and description of a specific concept, event, or person. Through the entry format, people can systematically understand and master the relevant knowledge in a certain field or concept, improving their overall understanding and grasp of knowledge.

[0068] In this embodiment, the structured comprehensive character knowledge entry is a knowledge entry with a clear structure and organization obtained after structured parsing of the unstructured comprehensive character knowledge entry. Taking the title of an unstructured article on healthy diet as an example, after structured parsing, a knowledge entry composed of structured sub-information bodies such as "food category-nutrient composition-recommended intake-applicable population" may be formed. The comprehensive character class represents a knowledge entry that may contain characters of multiple character categories, such as letters, numbers, special symbols, etc.

[0069] In this embodiment, all unstructured comprehensive character knowledge items in the knowledge base are structured and parsed to obtain all structured comprehensive character knowledge items. For example, for a large number of unstructured descriptive documents about historical events, by identifying key information therein, such as the name of the event, the time of occurrence, the location, the main characters, the course of events, etc., this information is organized into knowledge items with a clear structure. Just like parsing a vague description such as "In ancient times, an important war occurred, involving many people and a complicated process" into a structured knowledge item of "War name: [specific name], time of occurrence: [specific time], location: [specific location], main characters: [list of characters], course of events: [detailed process]".

[0070] In this embodiment, the vectorization model is a tool model that can convert structured, comprehensive character-based knowledge items into high-dimensional vectors. For example, the common word vector model Word2Vec can convert words into vectors to represent their semantics. In this system, the vectorization model similarly converts the various encodings of structured knowledge items into high-dimensional vectors, ensuring that each vector accurately represents the corresponding knowledge item.

[0071] In this embodiment, the high-dimensional vector is a vector composed of multiple dimensional numerical values, which is used to represent structured comprehensive character-type knowledge items in a high-dimensional space. These dimensions contain various features of the knowledge items (such as the features of the positions of the sub-information bodies contained therein) and semantic information. Assuming that a high-dimensional vector has 200 dimensions, some of the dimensions may represent the subject categories of the knowledge items, some of the dimensions may represent the strength of specific concepts, and some of the dimensions may represent the positional features of the sub-information bodies in the knowledge items. For example, for a knowledge item about "artificial intelligence algorithm" converted into a high-dimensional vector, one dimension may represent the degree to which the algorithm is a supervised learning algorithm, another dimension may represent its relevance in the field of image recognition, and another dimension may represent the positional features of the sub-information bodies it contains.

[0072] In this embodiment, the high-dimensional vector of the structured comprehensive character knowledge item is obtained by transforming the structured comprehensive character knowledge item using a vectorization model. It carries the detailed semantics and positional structure information of the corresponding knowledge item and is a digital representation of the knowledge item in vector space.

[0073] In this embodiment, user query vectors are generated by converting the query content entered by the user into a vector form using a specific method. For example, if a user enters "How to improve user activity on e-commerce platforms," the system will analyze this query, extract key information such as "e-commerce platform" and "user activity," and then convert this information into a high-dimensional vector using the same method used in the vectorization model.

[0074] In this embodiment, the multi-dimensional matching score is a score obtained by calculating the similarity between the user query vector and the high-dimensional vector of each comprehensive character knowledge item in the knowledge base in multiple dimensions.

[0075] In this embodiment, the user query domain knowledge graph is a knowledge network constructed based on the user query domain. It describes the relationships between various entities (such as concepts, things, etc.) in the domain, as well as the attributes and semantic information of the entities.

[0076] In this embodiment, a synonymous entity combination is a set of semantically similar entities in the knowledge graph of the domain to which the user query belongs. For example, in the knowledge graph of the education domain, "course" and "teaching content" may constitute a synonymous entity combination because they are semantically similar and both are related to education and teaching.

[0077] In this embodiment, the semantic and position constraint is an indicator that measures the degree to which synonymous entity combinations are restricted in terms of semantic similarity and positional identity. It is determined based on the ranking values of the synonymous entity combination in all the reference knowledge items to which it belongs, combined with factors such as comprehensive similarity and proximity. For example, in a knowledge graph about technological products, "smartphone" and "mobile phone" constitute a synonymous entity combination. If they always appear in the same position in most reference knowledge items (or the ranking values of the sub-information bodies in the knowledge items are very similar), and the semantic similarity is very high, then the semantic and position constraint of this synonymous entity combination is high, indicating that the relationship between them is close and fixed, and has a high importance in knowledge matching.

[0078] In this embodiment, the multi-dimensional weighting of the multi-dimensional matching score is to assign different weights to different dimensions of the multi-dimensional matching score according to the semantics and position constraints of the combination of synonymous entities.

[0079] In this embodiment, the matching result is a result of the matching degree between each structured comprehensive character knowledge item and the user query obtained through weighted summation based on multi-dimensional matching scores and multi-dimensional assigned weights.

[0080] In this embodiment, a personalized search report is generated based on the matching results, and provides the user with a report of knowledge related to the query.

[0081] The beneficial effects of the above technologies are: realizing effective calculation and matching of user queries and knowledge items in vector space, introducing knowledge graph of the field to which user queries belong, realizing accurate and effective weighting of multi-dimensional matching scores of user queries and knowledge items in vector space, thereby improving the matching retrieval accuracy between user queries and knowledge items, and providing high-quality retrieval results.

[0082] Example 2: Based on Example 1, the knowledge item vectorization module, refer to Figure 2 ,include:

[0083] The knowledge item encoding submodule is used to perform absolute position encoding and relative position association encoding on all sub-information bodies in each structured comprehensive character knowledge item, and generate the position and semantic perception vector of each sub-information body;

[0084] The feature alignment submodule is used to dimensionally align the positions and semantic perception vectors of all character class sub-information bodies in each structured comprehensive character class knowledge entry and map them to the same feature space to obtain a high-dimensional vector of the comprehensive character class knowledge entry for each structured comprehensive character class knowledge entry.

[0085] In this embodiment, sub-information bodies are the basic units that constitute structured comprehensive character knowledge items. Each structured comprehensive character knowledge item can be divided into multiple sub-information bodies, each of which carries part of the semantic content of the knowledge item. For example, a structured comprehensive character knowledge item about "What are the character types in a field?" may be divided into sub-information bodies such as "Field," "Character Type," and "What are there," which together constitute the complete knowledge content.

[0086] In this embodiment, the position- and semantics-aware vector is generated by encoding the absolute position and relative position association of each sub-information entity within each structured comprehensive character knowledge item. It not only incorporates the semantic information of the sub-information entity itself but also its position within the knowledge item, enabling the vector to more comprehensively represent the sub-information entity.

[0087] In this embodiment, the positions of the sub-information bodies and semantic perception vectors of all character classes in each structured comprehensive character class knowledge entry are dimensionally aligned and mapped to the same feature space to obtain a high-dimensional vector of the comprehensive character class knowledge entry of each structured comprehensive character class knowledge entry. That is, the maximum dimension of the positions of the sub-information bodies and the semantic perception vectors of all character classes is regarded as the dimension of the positions of the sub-information bodies and the semantic perception vectors of all character classes, and the empty element positions in the positions and semantic perception vectors after the dimension expansion are filled with 0.

[0088] The beneficial effect of these technologies is that they enable a more comprehensive integration of the structure and semantic information of knowledge items into vector representations, facilitating the system's precise understanding of the knowledge hierarchy and connotations. The feature alignment submodule aligns the position and semantic perception vectors of sub-information bodies and maps them to the same feature space, generating high-dimensional vectors for comprehensive character-based knowledge items. This addresses the issues of vector dimensionality and feature space inconsistency, improving the quality and consistency of knowledge vectorization, and thereby enhancing the efficiency and accuracy of knowledge matching and retrieval within the knowledge base system.

[0089] Example 3: Based on Example 2, the knowledge item encoding submodule, refer to Figure 2 ,include:

[0090] A basic coding unit is used to divide all character class sub-information bodies in each structured comprehensive character class knowledge entry, encode each sub-information body based on its semantic content, and obtain a basic code for each sub-information body;

[0091] An absolute position encoding unit, for generating an absolute position encoding of each sub-information body based on the different character class positions and the same character class positions of each sub-information body;

[0092] A relative position encoding unit, configured to generate a relative position association code for each sub-information body based on the relative association positions of the same character class and the relative association positions of the different character classes of each sub-information body;

[0093] The position and semantic perception vectorization unit is used to generate the position and semantic perception vector of each sub-information body based on the basic coding, absolute position coding, relative position coding and vectorization model of each sub-information body.

[0094] In this embodiment, the character class is a classification method for sub-information bodies in the structured comprehensive character class knowledge entry, for example, including digital class, letter class, special symbol class (such as "+", "!", etc.), etc.

[0095] In this embodiment, dividing the sub-information bodies of all character classes in each structured comprehensive character class knowledge entry means segmenting a complete structured comprehensive character class knowledge entry into specific sub-information bodies.

[0096] In this embodiment, encoding is performed based on the semantic content of each sub-information body to obtain a basic code for each sub-information body. This step is the process of giving the sub-information body an initial digital representation. This can be implemented based on an existing semantic encoder, such as the BERT encoding method.

[0097] In this embodiment, the basic code of the sub-information body is an initial digital representation obtained based on the semantic content coding of the sub-information body, which reflects the most basic semantic features of the sub-information body.

[0098] In this embodiment, the different character class position refers to the sorting position of each sub-information body relative to all other character class sub-information bodies in the corresponding structured comprehensive character class knowledge entry.

[0099] In this embodiment, the same character class position refers to the sorting position of each sub-information body in the set of sub-information bodies of the same character class to which it belongs.

[0100] In this embodiment, the absolute position code is generated based on the different character class positions and the same character class positions of each sub-information body. It combines the position of the sub-information body relative to all character class sub-information bodies in the entire knowledge entry, as well as the position information of the sub-information body in the set of same character class sub-information bodies.

[0101] In this embodiment, the relative associated position of the same character class is the sorting position of all synonymous sub-information bodies of each sub-information body of each structured comprehensive character class knowledge item in the same character class sub-information body set.

[0102] In this embodiment, the relative association position of different character classes is the sorting position of each sub-information body of each structured comprehensive character class knowledge entry in the sub-information body corresponding to all character classes and all synonymous sub-information bodies in the sub-information body corresponding to all character classes.

[0103] In this embodiment, the relative position association code is a code generated based on the relative association positions of the same character class and the relative association positions of different character classes of each sub-information body.

[0104] In this embodiment, the position and semantic perception vector of each sub-information body is generated based on the basic code, absolute position code, relative position code and vectorization model of each sub-information body. That is, the semantic basic code, absolute position code and relative position association code of the sub-information body are integrated, and then with the help of the vectorization model, a vector that comprehensively reflects the semantics and position information of the sub-information body is generated.

[0105] The beneficial effects of the above technologies are as follows: the basic coding unit encodes the semantic content of the sub-information body, lays the foundation for the understanding of knowledge, and can accurately capture the core meaning of the sub-information body. The absolute position coding unit uses the position of the sub-information body in different character classes and the same character class to generate absolute position coding, so that the system can clearly define the specific positioning of the sub-information body in the overall knowledge entry, which is helpful for sorting out the knowledge structure. The relative position coding unit generates relative position association coding based on the relative associated positions of the same and different character classes, explores the relationship between sub-information bodies, and further enriches the knowledge content. The position and semantic-aware vectorization unit combines the basic, absolute position, relative position coding and vectorization model to generate position and semantic-aware vectors, comprehensively integrates various types of information, and forms a vector representation rich in semantic and position information, which greatly improves the accuracy and completeness of knowledge representation, and provides a better data foundation for subsequent feature alignment and the efficient operation of the entire knowledge base system.

[0106] Example 4: Based on Example 3, the absolute position encoding unit, referring to Figure 2 ,include:

[0107] The variant character class positioning subunit is used to determine the sorting position of each sub-information body in the corresponding structured comprehensive character class knowledge entry, and generate the variant character class position of each sub-information body based on the sorting position of each sub-information body in the corresponding structured comprehensive character class knowledge entry;

[0108] The information body classification sub-unit is used to classify the sub-information bodies of the same character class in all the sub-information bodies of the character class in each structured comprehensive character class knowledge entry in an order-preserving manner to obtain a set of all sub-information bodies of the same character class in each structured comprehensive character class knowledge entry;

[0109] The same character class positioning subunit is used to determine the sorting position of each sub-information body in the set of sub-information bodies of the same character class to which it belongs, and generate the same character class position of each sub-information body based on the set of sub-information bodies of the same character class to which each sub-information body belongs and its sorting position in the set of sub-information bodies of the same character class to which it belongs;

[0110] The absolute position coding subunit is used to generate the absolute position coding of each sub-information body based on the different character class positions and the same character class positions of each sub-information body and the preset absolute position coding method.

[0111] In this embodiment, the sorting position of the sub-information body in the corresponding structured comprehensive character knowledge entry refers to the serial number of each sub-information body after being arranged in the original order in the structured comprehensive character knowledge entry. For example, taking the structured comprehensive character knowledge entry about "What are the character types of the field", the sorting values of the sub-information bodies "field", "character type" and "what are" contained therein are 1, 2 and 3 respectively.

[0112] In this embodiment, the different character class position of each sub-information body is generated based on the sorting position of each sub-information body in the corresponding structured comprehensive character class knowledge entry, that is, the sorting value of the sub-information body in the entire knowledge entry is regarded as its different character class position.

[0113] In this embodiment, the sub-information bodies of the same character class in all the sub-information bodies of the character class in each structured comprehensive character class knowledge entry are classified in an order-preserving manner to obtain a set of all sub-information bodies of the same character class in each structured comprehensive character class knowledge entry, which means that the sub-information bodies of the same character class in the same knowledge entry are grouped together and their order in the original knowledge entry is preserved.

[0114] In this embodiment, the set of sub-information bodies of the same character class is a set consisting of sub-information bodies of the same character class in the structured comprehensive character class knowledge entry, and the order of these sub-information bodies in the original knowledge entry is retained.

[0115] In this embodiment, the ranking position of a sub-information body in the set of sub-information bodies of the same character type refers to the ranking value of a sub-information body in the set of sub-information bodies of the same character type to which it belongs.

[0116] In this embodiment, the same-character-class position of each sub-information body is generated based on the set of sub-information bodies of the same character class to which each sub-information body belongs and its sorted position within the set of sub-information bodies of the same character class. This means that the same-character-class position is determined by combining the set of sub-information bodies of the same character class to which the sub-information body belongs and its sorted position within that set. For example, if the set of sub-information bodies of the same character class to which the sub-information body belongs is numbered 2 and its sorted position within that set is 5, then the same-character-class position of the sub-information body is a two-dimensional array (2, 5).

[0117] In this embodiment, the preset absolute position encoding method is a pre-set rule for converting the different character class positions and the same character class positions of the sub-information body into absolute position codes. For example, a simple encoding method can be set to multiply the different character class positions by 100 and add the same character class positions to obtain the absolute position code.

[0118] In this embodiment, based on the different character class position and the same character class position of each sub-information body and the preset absolute position encoding method, the absolute position code of each sub-information body is generated, that is, based on the different character class position and the same character class position determined previously, the calculation is performed according to the preset encoding method, thereby obtaining the absolute position code of each sub-information body. For example, in a knowledge entry, the "CPU" sub-information body under the "computer hardware" character class has a different character class position of 5 and a same character class position of 3. According to the preset absolute position encoding method (such as taking the different character class position as the hundreds place, the same character class position as the ones place, and filling the tens place with 0), the generated absolute position code is 503. This absolute position code integrates the position information of the sub-information body in the overall knowledge entry and the same character class set, which helps the system to accurately identify the unique position of the sub-information body in the knowledge structure.

[0119] The beneficial effects of the above technology are as follows: the different character class positioning sub-unit determines the sorting position of the sub-information body in the entire knowledge entry and generates the different character class position, so that the system can grasp the distribution of the sub-information body among different character classes from a macro level, and provide key positioning information for the overall layout of the knowledge entry. The information body classification sub-unit classifies the sub-information bodies of the same character class in an order-preserving manner, which not only retains the original order relationship between the sub-information bodies of the same character class, but also provides an ordered set for the subsequent determination of the position of the same character class, ensuring the integrity and coherence of the information of the same character class. The same character class positioning sub-unit further clarifies the sorting position of the sub-information body in the set of sub-information bodies of the same character class to which it belongs, generates the same character class position, and refines the positional relationship between the sub-information bodies of the same character class from a micro level, so that the system can more accurately identify the relative positions of the sub-information bodies of the same character class. The absolute position coding sub-unit generates absolute position coding by combining different character class positions, same character class positions and preset coding methods, and comprehensively integrates macro and micro position information, so that the absolute position coding of the sub-information body can reflect both its position in the entire knowledge entry and its position in the same character class set, greatly improving the accuracy and comprehensiveness of the position coding.

[0120] Example 5: Based on Example 3, the relative position encoding unit, referring to Figure 2 ,include:

[0121] The same-character-class relative association positioning subunit is used to treat the sorting position of all synonymous sub-information bodies of each sub-information body of each structured comprehensive character class knowledge item in the same-character-class sub-information body set as the same-character-class relative association position of each sub-information body;

[0122] The different character class relative association positioning subunit is used to use the sorting position of all synonymous sub-information bodies in the sub-information bodies corresponding to all character classes of each sub-information body of each structured comprehensive character class knowledge item as the different character class relative association position of each sub-information body;

[0123] The relative position coding subunit is used to generate a relative position association code for each sub-information body based on the relative associated positions of the same character class and the relative associated positions of different character classes of each sub-information body and a preset relative position association coding method.

[0124] In this embodiment, all synonymous sub-information bodies within the sub-information body set of the same character class to which a sub-information body belongs refer to other sub-information bodies within the sub-information body set of the same character class that have similar semantics to the specific sub-information body. For example, within the sub-information body set of the same character class "flower species," for the sub-information body "rose," "rose flower," "rose plant," and so on are synonymous sub-information bodies.

[0125] In this embodiment, the sorting position of all synonymous sub-information bodies within the set of sub-information bodies of the same character category refers to the sorting order of the sub-information bodies that are synonymous with the specific sub-information body within the set of sub-information bodies of the same character category. Continuing with the example of the set of sub-information bodies of the same character category "flower name," assuming the set order is ["rose," "rose," "rose," "lily"], for the sub-information body "rose," its synonymous sub-information body "rose" is sorted in position 1 within the set, and "rose" is sorted in position 3 within the set.

[0126] In this embodiment, the sorting position of the synonymous sub-information body in the sub-information bodies corresponding to all character classes refers to the arrangement order position of the synonymous sub-information body of a specific sub-information body in the set consisting of all character class sub-information bodies of the entire structured comprehensive character class knowledge entry.

[0127] In this embodiment, the preset relative position association encoding method is a predetermined rule for converting the relative position association positions of the same character class and the relative position association positions of different character classes in the sub-information body into relative position association codes. For example, a preset encoding method can be used to multiply the relative position association positions of the same character class by 10 and then add the relative position association positions of the different character classes to obtain the relative position association code.

[0128] In this embodiment, the relative position association code of each sub-information body is generated based on the relative association positions of the same character class and the relative association positions of different character classes of each sub-information body and the preset relative position association coding method. That is, according to the previously determined relative association positions of the same character class and the relative association positions of different character classes, the preset coding method is used to perform calculations to obtain the relative position association code of each sub-information body.

[0129] The beneficial effects of the above technology are: the relative associated position of the same character class can accurately depict the relative position relationship between sub-information bodies of the same character class based on semantic association, which helps the system understand the mutual position of sub-information bodies based on semantic similarity within the knowledge scope of the same character class, and explore the potential connection of knowledge of the same category. The relative associated position of different character classes enables the system to grasp the relative position of sub-information bodies based on semantic association from a more macro knowledge level, across different character classes, expand the dimension of knowledge association, and deepen the understanding of the semantic connection between different categories of knowledge. The relative position encoding subunit generates relative position association coding based on the above-mentioned relative associated positions of the same character class and different character classes and the preset encoding method, comprehensively integrates the semantic relative position information at different levels, and the generated coding contains rich semantic associated position features. This not only improves the degree of refinement of knowledge representation, enables knowledge vectors to more comprehensively and accurately reflect the semantic and positional relationships between knowledge, but also provides a more valuable data foundation for subsequent knowledge matching, retrieval and in-depth mining.

[0130] Example 6: Based on Example 1, the vector multi-dimensional matching module, refer to Figure 3 ,include:

[0131] The information body vectorization submodule is used to determine the semantic and position perception vectors of each sub-information body queried by the user based on the user query vector, and at the same time determine the semantic and position perception vectors of each sub-information body in each structured comprehensive character knowledge item based on the high-dimensional vector of each comprehensive character knowledge item in the knowledge base;

[0132] A semantic and position matching submodule is used to calculate the semantic and position matching score of each sub-information body queried by the user and each sub-information body in each structured comprehensive character knowledge entry based on the similarity of all elements in the same position in the semantic and position perception vector of each sub-information body queried by the user and the semantic and position perception vector of each sub-information body in each structured comprehensive character knowledge entry;

[0133] A matching score matrix submodule is used to construct a semantic and position matching matrix between the user query and each structured comprehensive character knowledge item based on the semantic and position matching scores of each sub-information body in the user query and each structured comprehensive character knowledge item;

[0134] The multidimensional matching submodule is used to construct a mapping curve based on all diagonal elements and all non-diagonal elements of the semantic and position matching matrix between the user query and each structured comprehensive character knowledge item, and calculate the multidimensional matching score between the user query vector and the high-dimensional vector of each comprehensive character knowledge item in the knowledge base.

[0135] In this embodiment, the semantic and location-aware vectors of each sub-information body queried by the user are determined based on the user query vector. This means that the query vector converted from the user input content is decomposed into the semantic and location-aware vectors corresponding to each sub-information body (i.e., the elements of the meaning and location information represented by each sub-information body are extracted from the query vector and re-aggregated to form the semantic and location-aware vectors of the sub-information body). For example, if a user queries "how to improve the battery life of a laptop computer," the system first converts it into a query vector, and then decomposes the query vector based on the mapping relationship between the vector and the knowledge structure to determine the semantic and location-aware vectors of sub-information bodies such as "laptop computer" and "battery life."

[0136] In this embodiment, the semantic and position perception vectors of each sub-information body in each structured comprehensive character knowledge item are determined based on the high-dimensional vector of each comprehensive character knowledge item in the knowledge base, which means that for the knowledge items in the knowledge base that have been converted into high-dimensional vectors, the semantic and position perception vectors of each sub-information body are also decomposed by the same decomposition method as above.

[0137] In this embodiment, based on the similarity of the semantic and position-aware vectors of each sub-information body queried by the user and the semantic and position-aware vectors of each sub-information body in each structured comprehensive character-class knowledge entry, the semantic and position matching scores of each sub-information body queried by the user and each sub-information body in each structured comprehensive character-class knowledge entry are calculated. Taking the "fruit nutrition" knowledge field as an example, the "apple nutrition" sub-information body in the user query vector and the "apple nutrition" sub-information body in a certain knowledge entry vector in the knowledge base are compared with the elements in the same position in their semantic and position-aware vectors. By comparing the similarity of the values of the two position elements (such as the size of the difference, the proportional relationship, etc.), the average of the similarity of all the elements in the same position is used as the semantic and position matching score of the two sub-information bodies, thereby measuring the degree of matching between them.

[0138] In this embodiment, based on the semantic and position matching scores of each sub-information body of the user query and each sub-information body in each structured comprehensive character knowledge item, a semantic and position matching matrix between the user query and each structured comprehensive character knowledge item is constructed. Assume that the user query has 3 sub-information bodies, and each knowledge item also has 2 corresponding sub-information bodies. With the user query sub-information bodies as rows and the knowledge item sub-information bodies as columns, the matching scores are filled in the corresponding positions to form a matrix with 3 rows and 2 columns. For example, the element in the third row and second column represents the semantic and position matching score between the third information body in the user query and the second information body in the knowledge item.

[0139] The beneficial effects of the above technology are as follows: the information body vectorization submodule can determine the semantic and position-aware vectors of each sub-information body based on the user query vector and the high-dimensional vectors of the comprehensive character-based knowledge items in the knowledge base. This provides a precise and detailed vector foundation for subsequent matching operations, enabling the system to understand and compare knowledge information from both semantic and positional dimensions. The semantic and position matching submodule calculates matching scores based on the proximity of elements in the same position. This approach effectively combines the semantic and positional information of the sub-information bodies, avoiding the limitations of considering only one factor, greatly improving matching accuracy and more accurately measuring the degree of fit between the user query and the sub-information bodies in the knowledge items. The matching score matrixization submodule provides an intuitive and organized data structure for subsequent comprehensive matching analysis, facilitating the system's overall understanding of matching trends. The multidimensional matching submodule constructs mapping curves based on the diagonal and off-diagonal elements of the matching matrix and calculates multidimensional matching scores. This fully exploits the rich information contained in the matrix and comprehensively evaluates the matching degree between the user query and the knowledge items from multiple dimensions. This provides a more comprehensive and in-depth reflection of the relationship between the two, improving the reliability and comprehensiveness of the matching results. Overall, the vector multidimensional matching module significantly improves the quality and effect of knowledge matching through a series of sophisticated operations.

[0140] Example 7: Based on Example 6, the multi-dimensional matching submodule, refer to Figure 3 ,include:

[0141] A first mapping curve generating unit is configured to sort all elements greater than a first threshold in a semantic and position matching matrix between the user query and each structured comprehensive character knowledge item from largest to smallest to obtain a first element sequence, and to generate a first mapping curve based on the row and column ordinal numbers of all elements in the first element sequence in the semantic and position matching matrix;

[0142] A second mapping curve generating unit is configured to sort all diagonal elements greater than a second threshold value in a semantic and position matching matrix between the user query and each structured comprehensive character knowledge item from largest to smallest to obtain a second element sequence, and generate a second mapping curve based on the row and column ordinal numbers of all elements in the second element sequence in the semantic and position matching matrix;

[0143] A third mapping curve generating unit is configured to sort all non-diagonal elements of the semantic and position matching matrix between the user query and each structured comprehensive character knowledge item that are greater than a third threshold value from largest to smallest to obtain a third element sequence, and generate a third mapping curve based on the row and column ordinal numbers of all elements in the third element sequence in the semantic and position matching matrix;

[0144] A multidimensional matching degree calculation unit is used to treat the matching degree between the first mapping curve and the first standard mapping curve, the matching degree between the second mapping curve and the second standard mapping curve, and the matching degree between the third mapping curve and the third standard mapping curve as the multidimensional matching scores of the user query vector and the high-dimensional vector of each comprehensive character knowledge item in the knowledge base.

[0145] In this embodiment, the first threshold is a pre-set numerical standard used to filter out elements with a value greater than this value from all elements in the semantic and position matching matrix. This serves to filter out elements with a relatively high degree of matching for further analysis. For example, the first threshold is set to 0.7.

[0146] In this embodiment, the first element sequence is formed by arranging all elements in the semantic and position matching matrix that are greater than the first threshold in descending order. For example, after filtering by the first threshold, the obtained element values are 0.85, 0.8, 0.75, etc. These elements are sorted from largest to smallest as [0.85, 0.8, 0.75]. This ordered sequence is the first element sequence.

[0147] In this embodiment, the first mapping curve is generated based on the row and column ordinal numbers of all elements in the first element sequence in the semantic and position matching matrix. This is to construct a curve using the row and column number information of each element in the first element sequence in the original semantic and position matching matrix. For example, the first element 0.85 in the first element sequence is located in the 2nd row and 3rd column of the matrix, and the second element 0.8 is located in the 1st row and 4th column of the matrix. By using these row and column ordinal numbers as coordinate points (such as (2,3), (1,4), etc.), and connecting these points according to a certain mathematical method (such as linear interpolation, spline interpolation, etc.), the first mapping curve is generated.

[0148] In this embodiment, the second threshold is also a pre-set value, but it is specifically used as a criterion for screening all diagonal elements in the semantic and position matching matrix. Similar to the first threshold, its purpose is to select diagonal elements with a high degree of matching for in-depth analysis. For example, the second threshold is set to 0.65.

[0149] In this embodiment, the second element sequence is formed by arranging all diagonal elements in the semantic and position matching matrix that are greater than the second threshold in descending order. For example, if the diagonal elements are 0.72, 0.68, 0.66, etc., the elements greater than the second threshold of 0.65 are sorted from largest to smallest as [0.72, 0.68], which is the second element sequence.

[0150] In this embodiment, a second mapping curve is generated based on the row and column ordinal numbers of all elements in the second element sequence in the semantic and position matching matrix. The principle is similar to that of generating the first mapping curve. According to the row and column position of each element in the second element sequence on the diagonal of the matrix (since it is a diagonal element, the row and column numbers are the same), such as the position of an element 0.72 in the matrix is the 3rd row and 3rd column, these row and column ordinal numbers are used as coordinate points (such as (3,3)), and then these points are connected by appropriate mathematical methods to generate a second mapping curve. Although the horizontal and vertical coordinates of the second mapping curve are the same and the slopes of the connected second mapping curves are the same, the length of the second mapping curve will vary because the row and column values of the filtered elements in the matrix are different.

[0151] In this embodiment, the third threshold is also a preset value used to filter out elements greater than this value from all off-diagonal elements in the semantic and position matching matrix. This serves to focus on elements that are not in key corresponding positions (off-diagonal positions) but still have a high degree of matching. For example, the third threshold is set to 0.6.

[0152] In this embodiment, the third element sequence is a sequence obtained by arranging all off-diagonal elements in the semantic and position matching matrix that are greater than the third threshold in descending order. For example, if the off-diagonal elements include 0.7, 0.63, and 0.58, the elements greater than the third threshold of 0.6 are sorted from largest to smallest as [0.7, 0.63], which is the third element sequence.

[0153] In this embodiment, a third mapping curve is generated based on the row and column ordinal numbers of all elements in the third element sequence in the semantic and position matching matrix. Similarly, the curve is constructed based on the row and column position information of the elements in the third element sequence in the matrix. For example, element 0.7 in the third element sequence is located in row 1, column 2 in the matrix, and element 0.63 is located in row 2, column 4 in the matrix. These row and column ordinal numbers are used as coordinate points (e.g., (1, 2), (2, 4), etc.), and appropriate mathematical methods are used to connect these points to generate the third mapping curve.

[0154] In this embodiment, the first (second, third) standard mapping curve is a pre-defined curve with specific characteristics, serving as a reference standard for measuring the degree of match between the first (second, third) mapping curve and the corresponding one. These standard mapping curves are determined based on the system's design objectives, the characteristics of the knowledge domain, and a large amount of experimental or empirical data. For example, within a specific knowledge domain, through analysis of a large amount of sample data, a curve representing the distribution characteristics of elements with a high degree of matching under an ideal matching state is determined as the first standard mapping curve. This curve provides a benchmark for evaluating the degree of closeness of the actual generated mapping curve to the ideal matching state.

[0155] In this embodiment, the degree of matching between the first mapping curve and the first standard mapping curve, the degree of matching between the second mapping curve and the second standard mapping curve, and the degree of matching between the third mapping curve and the third standard mapping curve are calculated by using a specific algorithm (for example, a curve similarity measurement algorithm can be used, such as calculating the Euclidean distance between two curves, cosine similarity, or other indicators to quantify their matching degree) to calculate the similarity between the actually generated mapping curve and the corresponding standard mapping curve.

[0156] The beneficial effects of the above technology include: by filtering and sorting the elements of the semantic and positional matching matrix according to different thresholds to generate mapping curves, it is possible to deeply analyze the matching matrix information from multiple perspectives. The first mapping curve is generated by sorting elements greater than a first threshold, which comprehensively grasps the distribution of highly matched elements and understands the prominent matching trends between user queries and knowledge items. The second mapping curve is generated by sorting diagonal elements greater than a second threshold, focusing on core position matching relationships and accurately assessing the fit between key corresponding positions. The third mapping curve is generated by sorting off-diagonal elements greater than a third threshold, which explores potential cross-position matching connections and enriches the matching analysis dimension. The matching degree between the generated mapping curve and the standard mapping curve is used as a multi-dimensional matching score, providing a scientific and quantitative evaluation standard. The actual matching degree is intuitively and accurately measured by comparing the actual results with the standard curve, making the score more convincing and reliable. This approach avoids the one-sidedness of single-dimensional analysis, comprehensively reflects complex semantic and positional relationships, improves the accuracy and depth of knowledge matching, provides users with accurate and detailed knowledge matching results, and optimizes the quality of knowledge retrieval services and user experience.

[0157] Example 8: Based on Example 1, the comprehensive weight matching module, refer to Figure 4 ,include:

[0158] The semantic and position constraint determination submodule is used to screen out all synonymous entity combinations in the knowledge graph of the domain to which the user queries belong, and determine the semantic and position constraint degrees of all synonymous entity combinations based on their ranking values in all related reference knowledge items;

[0159] A weighting submodule is used to weight the multidimensional matching scores based on the semantics and position constraints of all synonymous entity combinations to obtain multidimensional weighting of the multidimensional matching scores;

[0160] A calibration matching score generation submodule is used to obtain a calibration matching score between the user query and the corresponding structured comprehensive character knowledge item by weighted summation of the multi-dimensional matching scores based on the multi-dimensional weights assigned, and to use the calibration matching scores of all structured comprehensive character knowledge items in the knowledge base as matching results;

[0161] A first knowledge item screening submodule is configured to treat all structured comprehensive character knowledge items in the knowledge base whose calibration matching scores are greater than a first score threshold as all target knowledge items, and generate a semantic association graph between the user query and all structured comprehensive character knowledge items;

[0162] A second knowledge item screening submodule is configured to treat all structured comprehensive character knowledge items in the knowledge base whose calibration matching scores are greater than a second score threshold and not greater than the first score threshold as all expanded recommended knowledge items;

[0163] The personalized retrieval report generation submodule is used to generate a personalized retrieval report based on all target knowledge items and all expanded recommended knowledge items.

[0164] In this embodiment, filtering all synonymous entity combinations in the domain knowledge graph of the user's query refers to finding a set of semantically similar entities in the domain knowledge graph related to the user's query. For example, when a user queries for content related to the medical field, "hypertension" and "high blood pressure" may constitute a synonymous entity combination in the medical knowledge graph.

[0165] In this embodiment, the reference knowledge items refer to knowledge items used for reference and analysis of information related to the synonymous entity combination within the field to which the user query belongs. These knowledge items contain at least one entity (ie, sub-information body) in the synonymous entity combination.

[0166] In this embodiment, the ranking value of the synonymous entity combination in all the reference knowledge entries to which it belongs refers to the order in which the entities (or sub-information bodies) in the synonymous entity combination appear in the reference knowledge entries to which it belongs.

[0167] In this embodiment, based on the ranking values of all synonymous entity combinations in all their corresponding reference knowledge entries, the semantic and positional constraints of all synonymous entity combinations are determined, that is, by analyzing the similarity of the ranking values of two entities (or sub-information bodies) contained in the synonymous entity combination in the corresponding reference knowledge entries, combined with a certain calculation method, their semantic and positional constraints are determined.

[0168] In this embodiment, the multidimensional matching scores are weightedly added based on the multidimensional weights assigned to obtain the calibrated matching score of the user query and the corresponding structured comprehensive character knowledge entry, which means that according to the weights assigned to each dimension of the multidimensional matching score previously determined, the values of each dimension of the multidimensional matching score are multiplied by the corresponding weights and then added together to obtain a comprehensive calibrated matching score.

[0169] In this embodiment, the calibrated match score, derived through the weighted summation method described above, measures the degree of match between the user query and the corresponding structured comprehensive character knowledge item. This score combines the multi-dimensional match scores and the weighting information for each dimension, serving as an important basis for determining the relevance of the knowledge item to the user query.

[0170] In this embodiment, the first score threshold is a pre-set numerical standard used to filter out knowledge items that have a high degree of matching with the user query. For example, the first score threshold is set to 0.8.

[0171] In this embodiment, target knowledge items are structured, comprehensive character-based knowledge items in the knowledge base whose calibration match scores are greater than a first score threshold. These knowledge items have a high degree of match with the user's query and are considered to be the core knowledge content that best meets the user's needs. The system prioritizes presenting this type of knowledge to the user.

[0172] In this embodiment, the semantic association graph between the user query and all structured comprehensive character-based knowledge items is a graph that graphically displays the semantic relationship between the user query and all structured comprehensive character-based knowledge items in the knowledge base. In this graph, nodes can represent key entities in the user query and related entities in the knowledge items, and edges represent semantic connections between these entities, such as similarity relationships, causal relationships, etc. For example, if a user queries "application of artificial intelligence in image recognition", the graph will use "artificial intelligence", "image recognition", "application cases", etc. as nodes, and use edges to display their associations with the entities of each knowledge item in the knowledge base.

[0173] In this embodiment, the second score threshold is also a pre-set value that is smaller than the first score threshold. Its function is to filter out knowledge items that have a certain relevance to the user query but a slightly lower matching degree than the target knowledge item. For example, the second score threshold is set to 0.6.

[0174] In this embodiment, the expanded recommended knowledge items refer to structured comprehensive character knowledge items in the knowledge base whose calibration matching scores are greater than the second score threshold and not greater than the first score threshold.

[0175] In this embodiment, a personalized search report is generated based on all target knowledge items and all recommended expanded knowledge items. Specifically, the system organizes and sorts the selected target knowledge items and recommends them, presenting them to the user in a personalized manner. The report may include detailed information about the target knowledge items, a brief description of the recommended expanded knowledge items, and the aforementioned semantic association graph.

[0176] The beneficial effects of the above technologies are as follows: the semantic and position constraint determination submodule deeply explores the intrinsic connections between knowledge by screening synonymous entity combinations in the knowledge graph of the user's query field and determining their semantic and position constraints, providing an accurate basis for weight assignment. The weight assignment submodule assigns weights to multidimensional matching scores based on semantic and position constraints, making the weights scientific and reasonable, and highlighting the matching scores of key knowledge items. The calibration matching score generation submodule adds the multidimensional assigned weights and the multidimensional matching scores according to the weights to obtain the calibration matching score, which is used as the matching result, greatly improving the matching accuracy. The first and second knowledge item screening submodules use different score thresholds to hierarchically screen the target knowledge items and the expanded recommended knowledge items to meet the user's core and expanded knowledge needs. The personalized retrieval report generation submodule generates personalized retrieval reports based on the screened items, optimizes the user's retrieval experience, and improves the system's practicality and user satisfaction.

[0177] Example 9: Based on Example 8, the semantic and position constraint determination submodule is referenced. Figure 4 ,include:

[0178] The semantic and position constraint matrix generation submodule is used to filter out all synonymous entity combinations in the knowledge graph of the domain to which the user query belongs;

[0179] The semantic and position constraint calculation submodule is used to determine the semantic and position constraint of each synonymous entity combination based on the comprehensive similarity and proximity of the ranking values of each synonymous entity combination in all corresponding reference knowledge items.

[0180] In this embodiment, the comprehensive similarity of the ranking values of the synonymous entity combination in all the reference knowledge items to which it belongs is the ratio of the average ranking values of the two entities (sub-information bodies) included in the synonymous entity combination in all the reference knowledge items to which it belongs.

[0181] In this embodiment, the degree of similarity of a combination of similar entities mainly measures the degree of semantic similarity of each entity in the combination of similar entities and is obtained by searching a pre-prepared list containing the degrees of similarity of all combinations of similar entities.

[0182] In this embodiment, the semantic and positional constraints of each synonymous entity combination are determined based on the comprehensive similarity and proximity of the ranking values of each synonymous entity combination in all its corresponding reference knowledge items: that is, the product of the comprehensive similarity and proximity of the ranking values of each synonymous entity combination in all its corresponding reference knowledge items is used as the semantic and positional constraints of each synonymous entity combination.

[0183] The beneficial effect of the above technology is: by comprehensively considering the comprehensive similarity and proximity of the ranking values of each synonymous entity combination in all the reference knowledge items to which it belongs, the semantic and positional constraints are determined. This method comprehensively and meticulously mines the potential connections between entities in the knowledge items. The comprehensive similarity reflects the degree of similarity between different knowledge items, while the proximity highlights the semantic relevance of the synonymous entity combination itself. The combination of the two makes the determined semantic and positional constraints more accurate in reflecting the inherent logic between knowledge. This not only helps to improve the accuracy of understanding knowledge items, but also provides a more accurate weight basis for subsequent weight assignment sub-modules, thereby further optimizing the calculation of calibration matching scores, improving the accuracy and reliability of knowledge matching results, and ultimately providing users with knowledge retrieval results that better meet their needs and enhance the overall performance of the knowledge retrieval system.

[0184] Example 10: Based on Example 8, weights are assigned to submodules, refer to Figure 4 ,include:

[0185] a synonymous entity combination counting unit, configured to count the number of all synonymous entity combinations whose semantic and positional constraints are greater than a constraint threshold as a first number, and at the same time, count the number of all synonymous entity combinations whose semantic and positional constraints are not greater than the constraint threshold as a second number;

[0186] A multi-dimensional weight assigning unit, configured to obtain a multi-dimensional matching score based on the first quantity and the second quantity, comprising:

[0187] The weight of the matching degree between the first mapping curve and the first standard mapping curve is set to 0.5, and the product of the ratio of the first number to the total number of all synonymous entity combinations and 0.5 is used as the weight of the matching degree between the second mapping curve and the second standard mapping curve, and the product of the ratio of the second number to the total number of all synonymous entity combinations and 0.5 is used as the weight of the matching degree between the third mapping curve and the third standard mapping curve.

[0188] In this embodiment, the constraint degree threshold is a pre-set standard value used to classify or filter the semantic and positional constraints of synonymous entity combinations. For example, the constraint degree threshold is set to 0.7.

[0189] The beneficial effects of the above technology are: through such weight setting, the multi-dimensional matching score can more accurately reflect the matching degree between the user query and the structured comprehensive character knowledge items, effectively improve the accuracy of knowledge matching, and thus improve the quality of knowledge retrieval services, presenting users with knowledge content that better meets their needs, optimizing the user experience, and making the knowledge retrieval system more practical and reliable.

[0190] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the present invention and its equivalents, the present invention is intended to include these modifications and variations.

Claims

1. A knowledge base system that matches vectorized knowledge representation with large models, characterized by: include: A structured parsing module is used to perform structured parsing on all unstructured comprehensive character knowledge items in the knowledge base to obtain all structured comprehensive character knowledge items; A knowledge item vectorization module is used to convert each structured comprehensive character knowledge item into a high-dimensional vector based on a vectorization model, and obtain a comprehensive character knowledge item high-dimensional vector for each structured comprehensive character knowledge item; Vector multi-dimensional matching module, used to generate user query vectors and calculate the multi-dimensional matching scores between the user query vectors and the high-dimensional vectors of each comprehensive character knowledge item in the knowledge base; The comprehensive weight matching module is used to assign weights to multidimensional matching scores based on the semantics and position constraints of all synonymous entity combinations in the knowledge graph of the user's query field, obtain multidimensional weights of the multidimensional matching scores, obtain matching results based on the multidimensional matching scores and the multidimensional weights, and generate personalized retrieval reports based on the matching results.

2. The knowledge base system for matching vectorized knowledge representation with large models according to claim 1, characterized in that: Knowledge entry vectorization module, including: The knowledge item encoding submodule is used to perform absolute position encoding and relative position association encoding on all sub-information bodies in each structured comprehensive character knowledge item, and generate the position and semantic perception vector of each sub-information body; The feature alignment submodule is used to dimensionally align the positions and semantic perception vectors of all character class sub-information bodies in each structured comprehensive character class knowledge entry and map them to the same feature space to obtain a high-dimensional vector of the comprehensive character class knowledge entry for each structured comprehensive character class knowledge entry.

3. The knowledge base system for matching vectorized knowledge representation with large models according to claim 2, characterized in that: The knowledge item encoding submodule includes: A basic coding unit is used to divide all character class sub-information bodies in each structured comprehensive character class knowledge entry, encode each sub-information body based on its semantic content, and obtain a basic code for each sub-information body; An absolute position encoding unit, for generating an absolute position encoding of each sub-information body based on the different character class positions and the same character class positions of each sub-information body; A relative position encoding unit, configured to generate a relative position association code for each sub-information body based on the relative association positions of the same character class and the relative association positions of the different character classes of each sub-information body; The position and semantic perception vectorization unit is used to generate the position and semantic perception vector of each sub-information body based on the basic coding, absolute position coding, relative position coding and vectorization model of each sub-information body.

4. The knowledge base system for matching vectorized knowledge representation with large models according to claim 3, characterized in that: Absolute position encoding unit, including: The variant character class positioning subunit is used to determine the sorting position of each sub-information body in the corresponding structured comprehensive character class knowledge entry, and generate the variant character class position of each sub-information body based on the sorting position of each sub-information body in the corresponding structured comprehensive character class knowledge entry; The information body classification sub-unit is used to classify the sub-information bodies of the same character class in all the sub-information bodies of the character class in each structured comprehensive character class knowledge entry in an order-preserving manner to obtain a set of all sub-information bodies of the same character class in each structured comprehensive character class knowledge entry; The same character class positioning subunit is used to determine the sorting position of each sub-information body in the set of sub-information bodies of the same character class to which it belongs, and generate the same character class position of each sub-information body based on the set of sub-information bodies of the same character class to which each sub-information body belongs and its sorting position in the set of sub-information bodies of the same character class to which it belongs; The absolute position coding subunit is used to generate the absolute position coding of each sub-information body based on the different character class positions and the same character class positions of each sub-information body and the preset absolute position coding method.

5. The knowledge base system for matching vectorized knowledge representation with large models according to claim 3, characterized in that: Relative position encoding unit, including: The same-character-class relative association positioning subunit is used to treat the sorting position of all synonymous sub-information bodies of each sub-information body of each structured comprehensive character class knowledge item in the same-character-class sub-information body set as the same-character-class relative association position of each sub-information body; The different character class relative association positioning subunit is used to use the sorting position of all synonymous sub-information bodies in the sub-information bodies corresponding to all character classes of each sub-information body of each structured comprehensive character class knowledge item as the different character class relative association position of each sub-information body; The relative position coding subunit is used to generate a relative position association code for each sub-information body based on the relative associated positions of the same character class and the relative associated positions of different character classes of each sub-information body and a preset relative position association coding method.

6. The knowledge base system for matching vectorized knowledge representation with large models according to claim 1, characterized in that: Vector multi-dimensional matching module, including: The information body vectorization submodule is used to determine the semantic and position perception vectors of each sub-information body queried by the user based on the user query vector, and at the same time determine the semantic and position perception vectors of each sub-information body in each structured comprehensive character knowledge item based on the high-dimensional vector of each comprehensive character knowledge item in the knowledge base; A semantic and position matching submodule is used to calculate the semantic and position matching score of each sub-information body queried by the user and each sub-information body in each structured comprehensive character knowledge entry based on the similarity of all elements in the same position in the semantic and position perception vector of each sub-information body queried by the user and the semantic and position perception vector of each sub-information body in each structured comprehensive character knowledge entry; A matching score matrix submodule is used to construct a semantic and position matching matrix between the user query and each structured comprehensive character knowledge item based on the semantic and position matching scores of each sub-information body in the user query and each structured comprehensive character knowledge item; The multidimensional matching submodule is used to construct a mapping curve based on all diagonal elements and all non-diagonal elements of the semantic and position matching matrix between the user query and each structured comprehensive character knowledge item, and calculate the multidimensional matching score between the user query vector and the high-dimensional vector of each comprehensive character knowledge item in the knowledge base.

7. The knowledge base system for matching vectorized knowledge representation with large models according to claim 6, characterized in that: Multidimensional matching submodule, including: A first mapping curve generating unit is configured to sort all elements greater than a first threshold in a semantic and position matching matrix between the user query and each structured comprehensive character knowledge item from largest to smallest to obtain a first element sequence, and generate a first mapping curve based on the row and column ordinal numbers of all elements in the first element sequence in the semantic and position matching matrix; A second mapping curve generating unit is configured to sort all diagonal elements greater than a second threshold value in a semantic and position matching matrix between the user query and each structured comprehensive character knowledge item from largest to smallest to obtain a second element sequence, and generate a second mapping curve based on the row and column ordinal numbers of all elements in the second element sequence in the semantic and position matching matrix; A third mapping curve generating unit is configured to sort all non-diagonal elements of the semantic and position matching matrix between the user query and each structured comprehensive character knowledge item that are greater than a third threshold value from largest to smallest to obtain a third element sequence, and generate a third mapping curve based on the row and column ordinal numbers of all elements in the third element sequence in the semantic and position matching matrix; A multidimensional matching degree calculation unit is used to treat the matching degree between the first mapping curve and the first standard mapping curve, the matching degree between the second mapping curve and the second standard mapping curve, and the matching degree between the third mapping curve and the third standard mapping curve as the multidimensional matching scores of the user query vector and the high-dimensional vector of each comprehensive character knowledge item in the knowledge base.

8. The knowledge base system for matching vectorized knowledge representation with large models according to claim 1, characterized in that: Comprehensive weight matching module, including: The semantic and position constraint determination submodule is used to screen out all synonymous entity combinations in the knowledge graph of the domain to which the user queries belong, and determine the semantic and position constraint degrees of all synonymous entity combinations based on their ranking values in all related reference knowledge items; A weighting submodule is used to weight the multidimensional matching scores based on the semantics and position constraints of all synonymous entity combinations to obtain multidimensional weighting of the multidimensional matching scores; A calibration matching score generation submodule is used to obtain a calibration matching score between the user query and the corresponding structured comprehensive character knowledge item by weighted summation of the multi-dimensional matching scores based on the multi-dimensional weights assigned, and to use the calibration matching scores of all structured comprehensive character knowledge items in the knowledge base as matching results; A first knowledge item screening submodule is configured to treat all structured comprehensive character knowledge items in the knowledge base whose calibration matching scores are greater than a first score threshold as all target knowledge items, and generate a semantic association graph between the user query and all structured comprehensive character knowledge items; A second knowledge item screening submodule is configured to treat all structured comprehensive character knowledge items in the knowledge base whose calibration matching scores are greater than a second score threshold and not greater than the first score threshold as all expanded recommended knowledge items; The personalized retrieval report generation submodule is used to generate a personalized retrieval report based on all target knowledge items and all expanded recommended knowledge items.

9. The knowledge base system for matching vectorized knowledge representation with large models according to claim 8, characterized in that: Semantic and position constraint determination submodule includes: The semantic and position constraint matrix generation submodule is used to filter out all synonymous entity combinations in the knowledge graph of the domain to which the user query belongs; The semantic and position constraint calculation submodule is used to determine the semantic and position constraint of each synonymous entity combination based on the comprehensive similarity and proximity of the ranking values of each synonymous entity combination in all corresponding reference knowledge items.

10. The knowledge base system for matching vectorized knowledge representation with large models according to claim 8, characterized in that: Weights are assigned to submodules, including: a synonymous entity combination counting unit, configured to count the number of all synonymous entity combinations whose semantic and positional constraints are greater than a constraint threshold as a first number, and at the same time, count the number of all synonymous entity combinations whose semantic and positional constraints are not greater than the constraint threshold as a second number; A multi-dimensional weight assigning unit, configured to obtain a multi-dimensional matching score based on the first quantity and the second quantity, comprising: The weight of the matching degree between the first mapping curve and the first standard mapping curve is set to 0.5, and the product of the ratio of the first number to the total number of all synonymous entity combinations and 0.5 is used as the weight of the matching degree between the second mapping curve and the second standard mapping curve, and the product of the ratio of the second number to the total number of all synonymous entity combinations and 0.5 is used as the weight of the matching degree between the third mapping curve and the third standard mapping curve.

Citation Information

Patent Citations

  • Semantic representation method for patent text vectors

    CN104199809A

  • Hyperspectral curve matching method based on absorption peak characteristic

    CN105528580A

  • Spatial relationship semantic analysis method based on knowledge graph

    CN114564966A

  • Text-pedestrian retrieval method based on bounding box extraction and semantic consistency constraint

    CN116842212A

  • Standard substance and standard substance retrieval and sorting method and system based on search engine

    CN119127970A