Vectorized knowledge representation and knowledge base system matching to large models
By performing structured parsing and vectorization of the knowledge base, and combining it with domain knowledge graphs for multi-dimensional matching score weighting, the efficiency and accuracy issues of knowledge base system operations and matching in vector space are solved, achieving high-quality personalized knowledge retrieval.
Patent Information
- Application Number
- CN202510645253.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2045-05-20
AI Technical Summary
Existing knowledge base systems struggle to perform efficient computations and matching in vector spaces, and lack the ability to assign reasonable weights to matching scores based on domain knowledge graphs, thus affecting the accuracy of matching and retrieval between user queries and knowledge entries.
A structured parsing module is introduced to parse unstructured knowledge items. The knowledge items are transformed into high-dimensional vectors through a vectorization model. The weights of multi-dimensional matching scores are assigned based on the knowledge graph of the user's query domain, and a personalized search report is generated.
It improves the accuracy of matching and retrieval between user queries and knowledge entries, provides high-quality search results, and meets users' query needs in specific fields.
Smart Images

Figure CN120470033B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of knowledge matching retrieval, and in particular to a vectorized knowledge representation and large model matching knowledge base system. BACKGROUND
[0002] In the era of information explosion, knowledge base systems are crucial for the storage, management and retrieval of knowledge. With the rapid growth of data volume and the increasing complexity of user demand, traditional knowledge base systems face many challenges. Accurate and efficient retrieval of information matching user demand from massive knowledge has become a pressing problem. Vectorized knowledge representation and large model matching technology brings new opportunities for the optimization of knowledge base systems. Converting knowledge items into vector form can better utilize the mathematical properties of vector space for similarity calculation and matching, improving the accuracy and efficiency of retrieval. At the same time, combined with large models, it can fully leverage the advantages of large models in semantic understanding, knowledge reasoning, etc., and deeply mine the associations between knowledge, providing more intelligent and accurate knowledge services for users. This combined knowledge base system has broad application prospects in many fields such as intelligent question answering, information retrieval, intelligent decision-making, etc., and is expected to push knowledge management and utilization to a new height.
[0003] However, existing knowledge base systems have some defects. Knowledge is difficult to operate and match effectively in vector space. Lack of ability to assign reasonable weights to matching scores based on domain knowledge graphs affects the matching retrieval accuracy between user queries and knowledge items, making it difficult for users to quickly obtain accurate and demand-compliant knowledge information.
[0004] Therefore, the present application proposes a vectorized knowledge representation and large model matching knowledge base system. SUMMARY
[0005] The present application provides a vectorized knowledge representation and large model matching knowledge base system to realize effective operation and matching of user queries and knowledge items in vector space, introduces the domain knowledge graph of user queries, realizes accurate and effective weight assignment of multi-dimensional matching scores of user queries and knowledge items in vector space, and improves the matching retrieval accuracy between user queries and knowledge items, and provides high-quality retrieval results.
[0006] The present application provides a vectorized knowledge representation and large model matching knowledge base system, comprising:
[0007] A structured analysis module is configured to structure and analyze all unstructured comprehensive character class knowledge items in the knowledge base to obtain all structured comprehensive character class knowledge items.
[0008] a knowledge item vectorization module, configured to convert each structured comprehensive character class knowledge item into a high-dimensional vector based on a vectorization model, to obtain a comprehensive character class knowledge item high-dimensional vector of each structured comprehensive character class knowledge item;
[0009] a vector multi-dimensional matching module, configured to generate a user query vector, and to calculate a multi-dimensional matching score of the user query vector and the comprehensive character class knowledge item high-dimensional vector of each comprehensive character class knowledge item in the knowledge base;
[0010] a comprehensive weight matching module, configured to assign weights to the multi-dimensional matching scores based on semantics and position constraints of all synonymous entity combinations in a domain knowledge graph to which the user query belongs, to obtain multi-dimensional assigned weights of the multi-dimensional matching scores, to obtain a matching result based on the multi-dimensional matching scores and the multi-dimensional assigned weights, and to generate a personalized search report based on the matching result.
[0011] Preferably, the knowledge item vectorization module comprises:
[0012] a knowledge item encoding submodule, configured to perform absolute position encoding and relative position association encoding on all sub-information bodies in each structured comprehensive character class knowledge item, to generate a position and semantic perception vector of each sub-information body;
[0013] a feature alignment submodule, configured to perform dimension alignment and mapping to the same feature space on the position and semantic perception vectors of all character class sub-information bodies in each structured comprehensive character class knowledge item, to obtain a comprehensive character class knowledge item high-dimensional vector of each structured comprehensive character class knowledge item.
[0014] Preferably, the knowledge item encoding submodule comprises:
[0015] a basic encoding unit, configured to divide all character class sub-information bodies in each structured comprehensive character class knowledge item, and to encode based on semantic content of each sub-information body to obtain a basic encoding of each sub-information body;
[0016] an absolute position encoding unit, configured to generate an absolute position encoding of each sub-information body based on a different character class position and a same character class position of each sub-information body;
[0017] a relative position encoding unit, configured to generate a relative position association encoding of each sub-information body based on a same character class relative association position and a different character class relative association position of each sub-information body;
[0018] a position and semantic perception vectorization unit, configured to generate a position and semantic perception vector of each sub-information body based on the basic encoding, the absolute position encoding, the relative position encoding of each sub-information body, and a vectorization model.
[0019] Preferably, the absolute position encoding unit comprises:
[0020] The different character class positioning sub-unit is configured to determine the order position of each sub-information body in the corresponding structured comprehensive character class knowledge item, and generate the different character class position of each sub-information body based on the order position of each sub-information body in the corresponding structured comprehensive character class knowledge item.
[0021] The information body classification sub-unit is configured to sequentially and reservedly classify the same character class sub-information bodies in the sub-information bodies of each character class in each structured comprehensive character class knowledge item, and obtain a same character class sub-information body set of each structured comprehensive character class knowledge item.
[0022] The same character class positioning sub-unit is configured to determine the order position of each sub-information body in the same character class sub-information body set, and generate the same character class position of each sub-information body based on the same character class sub-information body set and the order position in the same character class sub-information body set.
[0023] The absolute position coding sub-unit is configured to generate the absolute position coding of each sub-information body based on the different character class position and the same character class position of each sub-information body and a preset absolute position coding mode.
[0024] Preferably, the relative position coding unit comprises:
[0025] The same character class relative correlation positioning sub-unit is configured to take the order positions of all synonymous sub-information bodies of each sub-information body of each structured comprehensive character class knowledge item in the same character class sub-information body set as the same character class relative correlation positions of each sub-information body.
[0026] The different character class relative correlation positioning sub-unit is configured to take the order positions of all synonymous sub-information bodies of each sub-information body of each structured comprehensive character class knowledge item in the corresponding sub-information bodies of all character classes as the different character class relative correlation positions of each sub-information body.
[0027] The relative position coding sub-unit is configured to generate the relative position correlation coding of each sub-information body based on the same character class relative correlation positions and the different character class relative correlation positions of each sub-information body and a preset relative position correlation coding mode.
[0028] Preferably, the vector multi-dimensional matching module comprises:
[0029] an information body vectorization submodule configured to determine semantic and location-aware vectors of each sub-information body of a user query based on a user query vector, and determine semantic and location-aware vectors of each sub-information body in each structured comprehensive character class knowledge item based on a high-dimensional vector of each structured comprehensive character class knowledge item in a knowledge base;
[0030] a semantic and location matching submodule configured to calculate semantic and location matching scores of each sub-information body of the user query and each sub-information body in each structured comprehensive character class knowledge item based on similarities of elements at all same positions in the semantic and location-aware vectors of each sub-information body of the user query and the semantic and location-aware vectors of each sub-information body in each structured comprehensive character class knowledge item;
[0031] a matching score matrixing submodule configured to construct semantic and location matching matrices of the user query and each structured comprehensive character class knowledge item based on the semantic and location matching scores of each sub-information body of the user query and each sub-information body in each structured comprehensive character class knowledge item;
[0032] a multidimensional matching submodule configured to construct mapping curves based on all diagonal elements and all non-diagonal elements of the semantic and location matching matrices of the user query and each structured comprehensive character class knowledge item, and calculate multidimensional matching scores of the user query vector and the high-dimensional vector of each structured comprehensive character class knowledge item in the knowledge base.
[0033] Preferably, the multidimensional matching submodule comprises:
[0034] a first mapping curve generating unit configured to sort all elements greater than a first threshold value in all elements of the semantic and location matching matrices of the user query and each structured comprehensive character class knowledge item in descending order to obtain a first element sequence, and generate a first mapping curve based on row and column ordinal numbers of all elements in the first element sequence in the semantic and location matching matrices;
[0035] a second mapping curve generating unit configured to sort all elements greater than a second threshold value in all diagonal elements of the semantic and location matching matrices of the user query and each structured comprehensive character class knowledge item in descending order to obtain a second element sequence, and generate a second mapping curve based on row and column ordinal numbers of all elements in the second element sequence in the semantic and location matching matrices;
[0036] a third mapping curve generating unit configured to sort all elements greater than a third threshold value in all non-diagonal elements of the semantic and location matching matrices of the user query and each structured comprehensive character class knowledge item in descending order to obtain a third element sequence, and generate a third mapping curve based on row and column ordinal numbers of all elements in the third element sequence in the semantic and location matching matrices.
[0037] The multi-dimensional matching degree calculation unit is configured to take the matching degrees of the first mapping curve and the first standard mapping curve, the matching degrees of the second mapping curve and the second standard mapping curve, and the matching degrees of the third mapping curve and the third standard mapping curve as multi-dimensional matching scores of the user query vector and high-dimensional vectors of each comprehensive character class knowledge item in the knowledge base.
[0038] Preferably, the comprehensive weight matching module comprises:
[0039] The semantic and position constraint degree determination submodule is configured to filter all synonym entity combinations in the domain knowledge graph to which the user query belongs, and determine semantic and position constraint degrees of all the synonym entity combinations based on ranking values of all the synonym entity combinations in all the reference knowledge items to which the user query belongs.
[0040] The weight assignment submodule is configured to assign weights to the multi-dimensional matching scores based on the semantic and position constraint degrees of all the synonym entity combinations to obtain multi-dimensional weight assignments of the multi-dimensional matching scores.
[0041] The calibrated matching score generation submodule is configured to sum the multi-dimensional matching scores by weight based on the multi-dimensional weight assignments to obtain calibrated matching scores of the user query and the structured comprehensive character class knowledge items, and take the calibrated matching scores of all the structured comprehensive character class knowledge items in the knowledge base as matching results.
[0042] The first knowledge item filtering submodule is configured to take all the structured comprehensive character class knowledge items in the knowledge base with calibrated matching scores greater than a first score threshold as all target knowledge items, and generate a semantic association graph of the user query and all the structured comprehensive character class knowledge items.
[0043] The second knowledge item filtering submodule is configured to take all the structured comprehensive character class knowledge items in the knowledge base with calibrated matching scores greater than a second score threshold and not greater than the first score threshold as all expanded recommended knowledge items.
[0044] The personalized search report generation submodule is configured to generate a personalized search report based on all the target knowledge items and all the expanded recommended knowledge items.
[0045] Preferably, the semantic and position constraint degree determination submodule comprises:
[0046] The semantic and position constraint matrix generation submodule is configured to filter all synonym entity combinations in the domain knowledge graph to which the user query belongs.
[0047] The semantic and position constraint degree calculation submodule is configured to determine semantic and position constraint degrees of each synonym entity combination based on comprehensive similarity and synonymity of ranking values of each synonym entity combination in all the reference knowledge items to which the user query belongs.
[0048] Preferably, the weight assignment sub-module comprises:
[0049] The near-synonym entity combination statistics unit is configured to count the number of all near-synonym entity combinations with a semantic and position constraint degree greater than a constraint degree threshold as a first number, and count the number of all near-synonym entity combinations with a semantic and position constraint degree not greater than the constraint degree threshold as a second number.
[0050] The multi-dimensional weight assignment unit is configured to obtain a multi-dimensional assignment weight of the multi-dimensional matching score based on the first number and the second number, comprising:
[0051] The weight of the matching degree of the first mapping curve and the first standard mapping curve is set to 0.5, the product of the ratio of the first number to the total number of all near-synonym entity combinations and 0.5 is taken as the weight of the matching degree of the second mapping curve and the second standard mapping curve, and the product of the ratio of the second number to the total number of all near-synonym entity combinations and 0.5 is taken as the weight of the matching degree of the third mapping curve and the third standard mapping curve.
[0052] The beneficial effects of the present application relative to the prior art are: the structured analysis module performs structured analysis on the unstructured comprehensive character class knowledge entries in the knowledge base, which organizes the chaotic knowledge and makes it more organized and easier to understand, greatly improving the processability and retrievability of the knowledge, facilitating the subsequent system to deeply mine and utilize the knowledge. The knowledge entry vectorization module converts the structured knowledge entries into high-dimensional vectors by means of the vectorization model, representing the knowledge in the form of vectors, which can more accurately depict the internal characteristics and semantic information of the knowledge, facilitating the computer to efficiently store, calculate and analyze, and providing a more accurate data basis for knowledge matching. The vector multi-dimensional matching module generates a user query vector and calculates the multi-dimensional matching score of the high-dimensional vector of the knowledge entry in the knowledge base, evaluating the matching degree of the knowledge and the user query from multiple dimensions, making the matching process more comprehensive and detailed, and improving the accuracy and rationality of the matching, avoiding the one-sidedness of single-dimensional matching. The comprehensive weight matching module assigns weights to the multi-dimensional matching scores based on the semantic and position constraint degrees of the near-synonym entity combinations in the domain knowledge graph, fully considering the semantic relationships and structural characteristics of the knowledge domain, further optimizing the matching results, and better meeting the user's query needs in a specific field, generating a personalized retrieval report that better meets the user's actual needs. Through the collaborative work of each module, the entire system realizes the efficient operation of the whole process from structured processing to vector representation, accurate matching and personalized report generation, providing high-quality and personalized knowledge retrieval services for users, and helping to improve the efficiency and effectiveness of knowledge acquisition and utilization, which has important value in the field of knowledge management and application.
[0053] Additional features and advantages of the present application will be set forth in the description that follows, and in part will be apparent from the description, or can be learned by practice of the application. The objectives and other advantages of the present application will be realized and attained by the structure particularly pointed out in the written description and claims hereof as well as the appended drawings.
[0054] The technical solutions of the present application are described in further detail below with the aid of the accompanying drawings and examples. BRIEF DESCRIPTION OF DRAWINGS
[0055] The accompanying drawings are included to provide a further understanding of the present application and are incorporated in and constitute a part of the specification, illustrate embodiments of the present application and are used to explain the present application, but do not constitute a limitation on the present application. In the drawings:
[0056] Figure 1 A vectorized knowledge representation and large model matching knowledge base system in an embodiment of the present application;
[0057] Figure 2 An architecture diagram of a knowledge entry vectorization module in an embodiment of the present application;
[0058] Figure 3 An architecture diagram of a vector multidimensional matching module in an embodiment of the present application;
[0059] Figure 4 An architecture diagram of a comprehensive weight matching module in an embodiment of the present application. DETAILED DESCRIPTION
[0060] The preferred embodiments of the present application are described below with reference to the accompanying drawings, and it should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and do not constitute a limitation on the present application.
[0061] Embodiment 1: The present application provides a vectorized knowledge representation and large model matching knowledge base system, referring to Figure 1 , including:
[0062] A structured analysis module for structurally analyzing all unstructured comprehensive character class knowledge entries in the knowledge base to obtain all structured comprehensive character class knowledge entries;
[0063] A knowledge entry vectorization module for converting each structured comprehensive character class knowledge entry into a high-dimensional vector based on a vectorization model to obtain a comprehensive character class knowledge entry high-dimensional vector of each structured comprehensive character class knowledge entry;
[0064] A vector multidimensional matching module for generating a user query vector and calculating a multidimensional matching score of the user query vector and the comprehensive character class knowledge entry high-dimensional vector of each comprehensive character class knowledge entry in the knowledge base;
[0065] The comprehensive weight matching module is configured to assign weights to the multi-dimensional matching scores based on semantics and position constraint degrees of all near-sense entity combinations in the domain knowledge graph to which the user query belongs, to obtain multi-dimensional weight assignment of the multi-dimensional matching scores, to obtain a matching result based on the multi-dimensional matching scores and the multi-dimensional weight assignment, and to generate a personalized search report based on the matching result.
[0066] In this embodiment, the knowledge base is a collection of stored knowledge that covers various types of knowledge information, just like a huge knowledge warehouse, providing data support for the entire knowledge matching and retrieval system. For example, a legal knowledge base may contain various legal regulations, legal case analysis, and other knowledge content. When a user performs a legal-related query, the system will search for matching information from this knowledge base.
[0067] In this embodiment, the unstructured comprehensive character class knowledge item refers to a series of independent items that have no specific structure, are presented in character form, and are relatively disorganized. These items can be a detailed analysis and elaboration of knowledge in a specific field, or a brief introduction and explanation of a specific concept, event, or person. Through the form of items, people can systematically understand and master the relevant knowledge of a certain field or concept, improving their overall understanding and grasp of the knowledge.
[0068] In this embodiment, the structured comprehensive character class knowledge item is obtained by structuring the unstructured comprehensive character class knowledge item, and has a clear structure and organization form. For example, a non-structured article title about healthy diet may form a knowledge item composed of structured sub-information bodies such as "food category-nutrient content-recommended intake-applicable population" after structured analysis. Comprehensive character class represents characters in the knowledge item that may include multiple character classes such as letters, numbers, and special symbols.
[0069] In this embodiment, all unstructured comprehensive character class knowledge items in the knowledge base are structured and analyzed to obtain all structured comprehensive character class knowledge items. For example, for a large number of non-structured description documents about historical events, key information such as event name, occurrence time, location, main characters, and event process is identified and organized into knowledge items with clear structure. For example, a vague description such as "In ancient times, there was an important war involving many characters and complex processes" is analyzed into a structured knowledge item "War name: [specific name], occurrence time: [specific time], location: [specific location], main characters: [list of characters], event process: [detailed process]".
[0070] In this embodiment, the vectorization model is a tool model that can convert structured comprehensive character class knowledge entries into high-dimensional vectors. For example, the common word vector model Word2Vec can convert words into vectors to represent their semantics. In this system, the vectorization model similarly converts various encodings of structured knowledge entries into high-dimensional vectors, so that each vector can accurately represent the corresponding knowledge entry.
[0071] In this embodiment, the high-dimensional vector is a vector composed of multiple dimension values, used to represent structured comprehensive character class knowledge entries in a high-dimensional space. These dimensions contain various features (such as features containing sub-information body positions) and semantic information of the knowledge entries. Assuming that a high-dimensional vector has 200 dimensions, some dimensions may represent the topic category of the knowledge entry, some dimensions represent the strength of a specific concept, and some dimensions represent the position characteristics of the sub-information bodies in the knowledge entry. For example, for a high-dimensional vector converted from a knowledge entry about "artificial intelligence algorithm", one dimension may represent the degree to which the algorithm is a supervised learning algorithm, another dimension may represent the relevance of its application to image recognition, and another dimension may represent the position characteristics of the sub-information bodies it contains.
[0072] In this embodiment, the high-dimensional vector of the structured comprehensive character class knowledge entry of the structured comprehensive character class knowledge entry is a high-dimensional vector obtained by converting the structured comprehensive character class knowledge entry through the vectorization model. It carries detailed semantic and position structure information of the corresponding knowledge entry, and is a digital representation of the knowledge entry in vector space.
[0073] In this embodiment, the user query vector is generated according to the user input query content, and a specific method is used to convert it into a vector form. For example, the user inputs "how to improve the user activity of the e-commerce platform", the system will analyze this query sentence and extract key information such as "e-commerce platform" and "user activity", and then convert these information into a high-dimensional vector through the same method as the vectorization model.
[0074] In this embodiment, the multi-dimensional matching score is a score obtained by calculating the similarity of the user query vector and the high-dimensional vector of each comprehensive character class knowledge entry in the knowledge base in multiple dimensions.
[0075] In this embodiment, the user query domain knowledge graph is a knowledge network constructed based on the domain of the user query. It describes the relationships between various entities (such as concepts, things, etc.) in the domain, as well as the attributes and semantic information of the entities.
[0076] In this embodiment, the near-synonymous entity combination is a set of entities that are semantically similar in the knowledge graph of the domain to which the user query belongs. For example, in the knowledge graph of the education domain, "course" and "teaching content" may constitute a near-synonymous entity combination because they are semantically similar and both relate to education and teaching.
[0077] In this embodiment, the semantic and positional constraint degree is an index that measures the degree of restriction of the near-synonymous entity combination in terms of semantic similarity and positional similarity. It is determined based on the ranking values of the near-synonymous entity combination in all reference knowledge entries, combined with factors such as comprehensive similarity and near-synonym degree. For example, in a knowledge graph about technology products, "smartphone" and "mobile phone" constitute a near-synonymous entity combination. If in most reference knowledge entries, they always appear in the same position (or the ranking values of the sub-information bodies in the knowledge entries are highly similar), and the semantic similarity is high, then the semantic and positional constraint degree of this near-synonymous entity combination is high, indicating that the relationship between them is close and fixed, and they have high importance in knowledge matching.
[0078] In this embodiment, the multi-dimensional weight assignment of the multi-dimensional matching score is to assign different weights to different dimensions of the multi-dimensional matching score according to the semantic and positional constraint degree of the near-synonymous entity combination.
[0079] In this embodiment, the matching result is the matching degree result of each structured comprehensive character class knowledge entry and the user query obtained after calculation based on the multi-dimensional matching score and the multi-dimensional weight assignment.
[0080] In this embodiment, the personalized retrieval report is generated according to the matching result, providing a report of knowledge related to the query for the user.
[0081] The beneficial effects of the above technology are: to realize the effective operation and matching of the user query and the knowledge entry in the vector space, to introduce the knowledge graph of the domain to which the user query belongs, to realize the accurate and effective weight assignment of the multi-dimensional matching score of the user query and the knowledge entry in the vector space, and to improve the matching and retrieval precision between the user query and the knowledge entry, and to provide high-quality retrieval results.
[0082] Embodiment 2: Based on embodiment 1, the knowledge entry vectorization module, according to Figure 2 , comprises:
[0083] The knowledge entry encoding submodule is used for performing absolute position encoding and relative position association encoding on all sub-information bodies in each structured comprehensive character class knowledge entry, and generating a position and semantic perception vector for each sub-information body.
[0084] The feature alignment submodule is configured to align dimensions of the position and semantic perception vectors of all sub-information bodies of each structured comprehensive character class knowledge item and map the position and semantic perception vectors to the same feature space to obtain a comprehensive character class knowledge item high-dimensional vector of each structured comprehensive character class knowledge item.
[0085] In this embodiment, the sub-information body is a basic unit constituting the structured comprehensive character class knowledge item. Each structured comprehensive character class knowledge item can be divided into multiple sub-information bodies, each of which carries part of the semantic content of the knowledge item. For example, a structured comprehensive character class knowledge item about "what are the character types of the field" can be divided into sub-information bodies such as "field", "character type", "what are they", and the like, which together constitute the complete knowledge content.
[0086] In this embodiment, the position and semantic perception vector is a vector generated by performing absolute position coding and relative position association coding on the sub-information bodies in each structured comprehensive character class knowledge item. It not only contains the semantic information of the sub-information body itself, but also incorporates the position information of the sub-information body in the knowledge item, so that the vector can more comprehensively represent the sub-information body.
[0087] In this embodiment, aligning dimensions of the position and semantic perception vectors of all sub-information bodies of each structured comprehensive character class knowledge item and mapping the position and semantic perception vectors to the same feature space to obtain a comprehensive character class knowledge item high-dimensional vector of each structured comprehensive character class knowledge item means that the maximum dimension in the position and semantic perception vectors of all sub-information bodies is taken as the dimension of the position and semantic perception vectors of all sub-information bodies, and the empty element positions in the dimensionally expanded position and semantic perception vectors are filled with 0.
[0088] The above technology has the beneficial effect that it enables the structure and semantic information of the knowledge item to be more comprehensively incorporated into the vector representation, facilitating the system's accurate understanding of the hierarchy and connotation of the knowledge. The feature alignment submodule aligns dimensions of the position and semantic perception vectors of the sub-information bodies and maps the position and semantic perception vectors to the same feature space to generate a comprehensive character class knowledge item high-dimensional vector, solving the problem of inconsistent vector dimensions and feature spaces, improving the quality and consistency of the knowledge vectorization, and further improving the efficiency and accuracy of the knowledge matching and retrieval of the knowledge base system.
[0089] Embodiment 3: Based on Embodiment 2, the knowledge item coding submodule, referring to Figure 2 , comprises:
[0090] The basic coding unit is configured to divide all sub-information bodies of each structured comprehensive character class knowledge item, encode each sub-information body based on the semantic content of the sub-information body, and obtain a basic code of each sub-information body.
[0091] an absolute position encoding unit configured to generate an absolute position encoding of each sub-information body based on the heterogeneous character class position and the homogeneous character class position of each sub-information body;
[0092] a relative position encoding unit configured to generate a relative position encoding of each sub-information body based on the homogeneous character class relative associated position and the heterogeneous character class relative associated position of each sub-information body;
[0093] a position and semantic perception vectorization unit configured to generate a position and semantic perception vector of each sub-information body based on the base encoding, the absolute position encoding, the relative position encoding of each sub-information body, and a vectorization model.
[0094] In this embodiment, the character class is a classification method for sub-information bodies in a structured comprehensive character class knowledge item, for example, including a number class, a letter class, a special symbol class (such as “+”, “!” and the like), and the like.
[0095] In this embodiment, dividing all character class sub-information bodies in each structured comprehensive character class knowledge item is to divide a complete structured comprehensive character class knowledge item into specific sub-information bodies.
[0096] In this embodiment, the base encoding of each sub-information body is obtained based on the semantic content of each sub-information body, which is the process of giving the sub-information body an initial digital representation. It can be implemented based on existing semantic encoders, for example, based on the BERT encoding method.
[0097] In this embodiment, the base encoding of the sub-information body is an initial digital representation obtained based on the semantic content of the sub-information body, which reflects the most basic semantic features of the sub-information body.
[0098] In this embodiment, the heterogeneous character class position refers to the ordering position of each sub-information body in the corresponding structured comprehensive character class knowledge item relative to all other character class sub-information bodies.
[0099] In this embodiment, the homogeneous character class position refers to the ordering position of each sub-information body in the homogeneous character class sub-information body set.
[0100] In this embodiment, the absolute position encoding is an encoding generated based on the heterogeneous character class position and the homogeneous character class position of each sub-information body. It integrates the position of the sub-information body relative to all character class sub-information bodies in the entire knowledge item and the position information in the homogeneous character class sub-information body set.
[0101] In this embodiment, the homogeneous character class relative associated position is the ordering position of all synonymous sub-information bodies in the homogeneous character class sub-information body set in the homogeneous character class sub-information body set of each sub-information body of each structured comprehensive character class knowledge item.
[0102] In this embodiment, the relative position correlation position of different character classes is the ordering position of all synonymous sub-information bodies in the corresponding sub-information bodies of all character classes in each sub-information body of each structured comprehensive character class knowledge item.
[0103] In this embodiment, the relative position correlation coding is coding generated based on the relative position correlation position of the same character classes and the relative position correlation position of the different character classes of each sub-information body.
[0104] In this embodiment, the position and semantic perception vector of each sub-information body is generated based on the basic coding, the absolute position coding, the relative position coding of each sub-information body, and the vectorization model, that is, the semantic basic coding, the absolute position coding, and the relative position correlation coding of the sub-information body are integrated, and then the vectorization model is used to generate a vector that comprehensively reflects the semantic and position information of the sub-information body.
[0105] The beneficial effects of the above technology are as follows: the basic coding unit encodes the semantic content of the sub-information body, lays the foundation for understanding the knowledge, and can accurately capture the core meaning of the sub-information body. The absolute position coding unit generates the absolute position coding by using the position of the sub-information body in different character classes and the same character class, so that the system can clearly position the sub-information body in the overall knowledge item, which helps to sort out the knowledge structure. The relative position coding unit generates the relative position correlation coding based on the relative position correlation position of the same and different character classes, which excavates the mutual relationship between the sub-information bodies and further enriches the knowledge connotation. The position and semantic perception vectorization unit generates the position and semantic perception vector by combining the basic, absolute position, and relative position coding and the vectorization model, comprehensively integrates various types of information, forms a vector representation rich in semantic and position information, greatly improves the accuracy and integrity of the knowledge representation, and provides a better data basis for subsequent feature alignment and efficient operation of the entire knowledge base system.
[0106] Embodiment 4: Based on Embodiment 3, the absolute position coding unit, referring to Figure 2 , comprises:
[0107] The different character class positioning sub-unit is configured to determine the ordering position of each sub-information body in the corresponding structured comprehensive character class knowledge item, and generate the different character class position of each sub-information body based on the ordering position of each sub-information body in the corresponding structured comprehensive character class knowledge item.
[0108] The information body classification sub-unit is configured to sequentially preserve and classify the same character class sub-information bodies in the sub-information bodies of all character classes in each structured comprehensive character class knowledge item, and obtain a set of all same character class sub-information bodies of each structured comprehensive character class knowledge item.
[0109] The same character class positioning unit is configured to determine the order position of each sub-information body in the same character class sub-information body set to which the sub-information body belongs, and generate the same character class position of each sub-information body based on the same character class sub-information body set to which the sub-information body belongs and the order position in the same character class sub-information body set.
[0110] The absolute position encoding sub-unit is configured to generate the absolute position encoding of each sub-information body based on the different character class position and the same character class position of each sub-information body and a preset absolute position encoding mode.
[0111] In this embodiment, the order position of a sub-information body in a corresponding structured comprehensive character class knowledge entry refers to the serial number of the sub-information body after the sub-information body is arranged in the original order in the structured comprehensive character class knowledge entry. For example, the order values of the sub-information bodies "field", "character type" and "what" in the structured comprehensive character class knowledge entry "what are the character types of the field" are 1, 2 and 3, respectively.
[0112] In this embodiment, the different character class position of each sub-information body is generated based on the order position of the sub-information body in the corresponding structured comprehensive character class knowledge entry, that is, the order value of the sub-information body in the entire knowledge entry is taken as the different character class position of the sub-information body.
[0113] In this embodiment, the same character class sub-information bodies in all character class sub-information bodies in each structured comprehensive character class knowledge entry are classified in a sequence-preserving manner to obtain all same character class sub-information body sets of each structured comprehensive character class knowledge entry, that is, the same character class sub-information bodies in the same knowledge entry are classified together, and the order of the same character class sub-information bodies in the original knowledge entry is preserved.
[0114] In this embodiment, the same character class sub-information body set is a set composed of sub-information bodies of the same character class in the structured comprehensive character class knowledge entry, and the order of the sub-information bodies in the original knowledge entry is preserved.
[0115] In this embodiment, the order position of a sub-information body in the same character class sub-information body set to which the sub-information body belongs refers to the order value of the sub-information body in the same character class sub-information body set to which the sub-information body belongs.
[0116] In this embodiment, the homograph class position of each sub-information body is generated based on the homograph class sub-information body set to which the sub-information body belongs and the rank position in the homograph class sub-information body set, that is, the homograph class position is determined in combination with the homograph class sub-information body set in which the sub-information body is located and the rank position of the sub-information body in the set. For example, if the homograph class sub-information body set in which the sub-information body is located is numbered 2 and the rank position of the sub-information body in the set is 5, the homograph class position of the sub-information body is a two-dimensional array (2, 5).
[0117] In this embodiment, the preset absolute position encoding mode is a rule that is set in advance and is used to convert the heterograph class position and the homograph class position of the sub-information body into the absolute position encoding. For example, a simple encoding mode can be set, in which the heterograph class position is multiplied by 100 and then the homograph class position is added to obtain the absolute position encoding.
[0118] In this embodiment, the absolute position encoding of each sub-information body is generated based on the heterograph class position and the homograph class position of each sub-information body and the preset absolute position encoding mode, that is, the absolute position encoding of each sub-information body is obtained by calculating according to the preset encoding mode based on the heterograph class position and the homograph class position determined in the foregoing. For example, in a knowledge entry, the heterograph class position of the "CPU" sub-information body under the "computer hardware" character class is 5 and the homograph class position is 3, and the absolute position encoding generated according to the preset absolute position encoding mode (for example, the heterograph class position is taken as the hundred place, the homograph class position is taken as the unit place, and the ten place is filled with 0) is 503. This absolute position encoding integrates the position information of the sub-information body in the overall knowledge entry and the homograph class set, which is helpful for the system to accurately identify the unique position of the sub-information body in the knowledge structure.
[0119] The beneficial effects of the above technology are: the different character class positioning subunit determines the ordering position of the sub-information body in the entire knowledge entry and generates a different character class position, enabling the system to grasp the distribution of the sub-information body among different character classes from a macro perspective, and providing key positioning information for the overall layout of the knowledge entry. The information body classification subunit classifies the same character class sub-information body in a sequence-preserving manner, which not only preserves the original sequence relationship among the same character class sub-information bodies, but also provides an ordered set for subsequent determination of the same character class position, ensuring the integrity and coherence of the same character class information. The same character class positioning subunit further specifies the ordering position of the sub-information body in the same character class sub-information body set to which it belongs, and generates a same character class position, which refines the position relationship among the same character class sub-information bodies from a micro perspective, facilitating the system to more accurately identify the relative position of the same character class sub-information body. The absolute position encoding subunit generates an absolute position encoding by combining the different character class position, the same character class position, and the preset encoding mode, comprehensively integrating macro and micro position information, so that the absolute position encoding of the sub-information body can not only reflect its position in the entire knowledge entry, but also reflect its position in the same character class set, greatly improving the accuracy and comprehensiveness of the position encoding.
[0120] In embodiment 5, on the basis of embodiment 3, the relative position encoding unit, referring to Figure 2 , comprises:
[0121] The same character class relative association positioning subunit is configured to take the ordering position of all synonymous sub-information bodies of each sub-information body of each structured comprehensive character class knowledge entry in the same character class sub-information body set to which the synonymous sub-information bodies belong as the same character class relative association position of each sub-information body.
[0122] The different character class relative association positioning subunit is configured to take the ordering position of all synonymous sub-information bodies of each sub-information body of each structured comprehensive character class knowledge entry in the corresponding sub-information bodies of all character classes as the different character class relative association position of each sub-information body.
[0123] The relative position encoding subunit is configured to generate the relative position association encoding of each sub-information body based on the same character class relative association position and the different character class relative association position of each sub-information body and a preset relative position association encoding mode.
[0124] In this embodiment, the all synonymous sub-information bodies of a sub-information body in the same character class sub-information body set refer to other sub-information bodies with similar semantics to the specific sub-information body in the same character class sub-information body set. For example, in the same character class sub-information body set of "flower types", "rose flower" and "rose plant" are synonymous sub-information bodies of the sub-information body "rose".
[0125] In this embodiment, the ordering position of the synonymous sub-information body in the sub-information body set of the same character class to which the synonymous sub-information body belongs refers to the arrangement order position of the synonymous sub-information body of the specific sub-information body in the sub-information body set of the same character class. Taking the sub-information body set of the same character class of "flower name" as an example, assuming that the set order is ["rose", "roses", "Chinese rose", "lily"], for the sub-information body "roses", the ordering position of its synonymous sub-information body "rose" in the set is 1, and the ordering position of "Chinese rose" in the set is 3.
[0126] In this embodiment, the ordering position of the synonymous sub-information body in the sub-information body of all character classes refers to the arrangement order position of the synonymous sub-information body of the specific sub-information body in the set composed of all character class sub-information bodies of the structured comprehensive character class knowledge item.
[0127] In this embodiment, the preset relative position association coding mode is a rule determined in advance, which is used to convert the same character class relative association position and the different character class relative association position of the sub-information body into the relative position association coding. For example, a coding mode can be preset to multiply the same character class relative association position by 10, and then add the different character class relative association position to obtain the relative position association coding.
[0128] In this embodiment, based on the same character class relative association position and the different character class relative association position of each sub-information body and the preset relative position association coding mode, the relative position association coding of each sub-information body is generated, that is, according to the same character class relative association position and the different character class relative association position determined in advance, the relative position association coding of each sub-information body is obtained by performing operation according to the preset coding mode.
[0129] The beneficial effects of the above technology are: the same character class relative association position can accurately depict the relative position relationship between the same character class sub-information bodies based on semantic association, which helps the system to understand the mutual position of the sub-information bodies based on semantic similarity within the same character class knowledge, and to mine the potential relationship of the same class knowledge. The different character class relative association position enables the system to grasp the relative position of the sub-information bodies based on semantic association from a more macro knowledge level, across different character classes, expands the dimension of knowledge association, and deepens the cognition of the semantic relationship between different categories of knowledge. The relative position coding sub-unit generates the relative position association coding based on the above same character class and different character class relative association positions and the preset coding mode, which comprehensively integrates the semantic relative position information at different levels, and the generated coding contains rich semantic association position features. This not only improves the fine degree of knowledge representation, so that the knowledge vector can more comprehensively and accurately reflect the semantic and position relationship between knowledge, but also provides a more valuable data basis for subsequent knowledge matching, retrieval and deep mining.
[0130] Embodiment 6: Based on Embodiment 1, the vector multidimensional matching module, referring to Figure 3 , comprises:
[0131] an information body vectorization submodule, configured to determine semantic and location-aware vectors of each information body of a user query based on a query vector of the user query, and determine semantic and location-aware vectors of each information body in each structured comprehensive character class knowledge item based on a high-dimensional vector of each comprehensive character class knowledge item in a knowledge base;
[0132] a semantic and location matching submodule, configured to calculate semantic and location matching scores of each information body of the user query and each information body in each structured comprehensive character class knowledge item based on similarities of elements at all same positions in the semantic and location-aware vectors of each information body of the user query and the semantic and location-aware vectors of each information body in each structured comprehensive character class knowledge item;
[0133] a matching score matrixing submodule, configured to construct semantic and location matching matrices of the user query and each structured comprehensive character class knowledge item based on the semantic and location matching scores of each information body of the user query and each information body in each structured comprehensive character class knowledge item;
[0134] a multidimensional matching submodule, configured to construct mapping curves based on all diagonal elements and all non-diagonal elements of the semantic and location matching matrices of the user query and each structured comprehensive character class knowledge item, and calculate multidimensional matching scores of the query vector and the high-dimensional vector of each comprehensive character class knowledge item in the knowledge base.
[0135] In this embodiment, the semantic and location-aware vectors of each information body of the user query are determined based on the query vector converted from the user input content, and the semantic and location-aware vectors corresponding to each information body are extracted from the query vector (i.e., the elements representing the semantic and location information of a single information body in the query vector are extracted and reorganized to form the semantic and location-aware vectors of the information body). For example, the user query “how to improve the battery endurance of a notebook computer” is first converted into a query vector, and then the query vector is disassembled according to the mapping relationship between the vector and the knowledge structure to determine the semantic and location-aware vectors of information bodies such as “notebook computer” and “battery endurance”.
[0136] In this embodiment, the semantic and position perception vectors of each sub-information body in each structured comprehensive character class knowledge item are determined based on the high-dimensional vector of each comprehensive character class knowledge item in the knowledge base, which means that for the knowledge items in the knowledge base that have been converted into high-dimensional vectors, the semantic and position perception vectors of each sub-information body are also decomposed by the same decomposition method as described above.
[0137] In this embodiment, the semantic and position matching scores of each sub-information body of the user query and each sub-information body in each structured comprehensive character class knowledge item are calculated based on the semantic and position perception vectors of each sub-information body of the user query and each sub-information body in each structured comprehensive character class knowledge item. For example, in the knowledge field of "fruit nutrition", the "apple nutrition" sub-information body in the user query vector and the "apple nutrition" sub-information body in the vector of a certain knowledge item in the knowledge base are compared in terms of the same position elements in their semantic and position perception vectors. By comparing the similarity of the numerical values of the two position elements (such as the difference size, proportional relationship, etc.), the average of the similarity of all the same position elements is taken as the semantic and position matching score of the two sub-information bodies, which is used to measure the matching degree between them.
[0138] In this embodiment, the semantic and position matching matrix of the user query and each structured comprehensive character class knowledge item is constructed based on the semantic and position matching scores of each sub-information body of the user query and each sub-information body in each structured comprehensive character class knowledge item. Assuming that the user query has 3 sub-information bodies and each knowledge item has 2 corresponding sub-information bodies. Taking the user query sub-information bodies as rows and the knowledge item sub-information bodies as columns, the matching scores are filled in the corresponding positions to form a 3x2 matrix. For example, the element in the third row and second column represents the semantic and position matching score of the third information body in the user query and the second information body in the knowledge item.
[0139] The beneficial effects of the above technology are that the information body vectorization submodule can determine the semantic and location awareness vectors of each sub-information body based on the user query vector and the high-dimensional vector of the comprehensive character class knowledge item in the knowledge base, thereby providing a precise and detailed vector basis for subsequent matching operations, enabling the system to understand and compare knowledge information from the semantic and location dual dimensions. The semantic and location matching submodule calculates the matching score based on the proximity of the same position elements, which effectively combines the semantic and location information of the sub-information body, avoids the limitations of considering only a single factor, greatly improves the accuracy of matching, and can more accurately measure the degree of fit between the user query and the sub-information body in the knowledge item. The matching score matrix submodule provides an intuitive and ordered data structure for subsequent comprehensive analysis of the matching situation, making it easy for the system to grasp the matching situation from the overall level. The multi-dimensional matching submodule constructs a mapping curve based on the diagonal and non-diagonal elements of the matching matrix and calculates the multi-dimensional matching score, fully excavates the rich information contained in the matrix, comprehensively evaluates the matching degree of the user query and the knowledge item from multiple dimensions, and can more comprehensively and deeply reflect the relationship between the two, thereby improving the reliability and comprehensiveness of the matching result. Overall, the vector multi-dimensional matching module significantly improves the quality and effectiveness of knowledge matching through a series of fine operations.
[0140] In the embodiment 7, the multi-dimensional matching submodule, with reference to Figure 3 , comprises:
[0141] The first mapping curve generation unit is configured to sort all elements greater than the first threshold value in the semantic and location matching matrix of the user query and each structured comprehensive character class knowledge item from large to small to obtain a first element sequence, and generate a first mapping curve based on the row and column ordinal numbers of all elements in the semantic and location matching matrix in the first element sequence.
[0142] The second mapping curve generation unit is configured to sort all elements greater than the second threshold value in the diagonal elements of the semantic and location matching matrix of the user query and each structured comprehensive character class knowledge item from large to small to obtain a second element sequence, and generate a second mapping curve based on the row and column ordinal numbers of all elements in the semantic and location matching matrix in the second element sequence.
[0143] The third mapping curve generation unit is configured to sort all elements greater than the third threshold value in the non-diagonal elements of the semantic and location matching matrix of the user query and each structured comprehensive character class knowledge item from large to small to obtain a third element sequence, and generate a third mapping curve based on the row and column ordinal numbers of all elements in the semantic and location matching matrix in the third element sequence.
[0144] The multi-dimensional matching degree calculation unit is configured to take the matching degrees of the first mapping curve and the first standard mapping curve, the matching degrees of the second mapping curve and the second standard mapping curve, and the matching degrees of the third mapping curve and the third standard mapping curve as multi-dimensional matching scores of the user query vector and the high-dimensional vector of each comprehensive character class knowledge item in the knowledge base.
[0145] In this embodiment, the first threshold is a preset numerical standard, which is used to screen elements greater than the value in all elements of the semantic and position matching matrix. Its function is to screen out elements with relatively high matching degrees for further analysis. For example, the first threshold is set to 0.7.
[0146] In this embodiment, the first element sequence is a sequence formed by arranging all elements greater than the first threshold in the semantic and position matching matrix in descending order. For example, after screening by the first threshold, the element values obtained are 0.85, 0.8, 0.75, etc. Arranging these elements in descending order as [0.85, 0.8, 0.75] forms the first element sequence.
[0147] In this embodiment, the first mapping curve is generated based on the row and column ordinal numbers of all elements in the semantic and position matching matrix in the first element sequence. That is, the row number and column number of each element in the first element sequence in the original semantic and position matching matrix are used to construct a curve. For example, the first element 0.85 in the first element sequence is located at the 2nd row and the 3rd column in the matrix, and the second element 0.8 is located at the 1st row and the 4th column in the matrix. By taking these row and column ordinal numbers as coordinate points (such as (2, 3), (1, 4), etc.), and connecting these points according to a certain mathematical method (such as linear interpolation, spline interpolation, etc.), the first mapping curve is generated.
[0148] In this embodiment, the second threshold is also a preset numerical value, but it is a standard specifically used to screen all diagonal elements in the semantic and position matching matrix. Similar to the first threshold, its purpose is to select the part with a higher matching degree in the diagonal elements for in-depth analysis. For example, the second threshold is set to 0.65.
[0149] In this embodiment, the second element sequence is a sequence formed by arranging all diagonal elements greater than the second threshold in the semantic and position matching matrix in descending order. For example, the diagonal elements are 0.72, 0.68, 0.66, etc. The elements greater than the second threshold 0.65 are arranged in descending order as [0.72, 0.68], which is the second element sequence.
[0150] In this embodiment, the second mapping curve is generated based on the row and column ordinal numbers of all elements in the second element sequence in the semantic and position matching matrix. The principle is similar to that of generating the first mapping curve. According to the row and column positions of each element in the second element sequence on the diagonal line of the matrix (since it is a diagonal element, the row and column numbers are the same), such as the position of an element 0.72 in the matrix is the 3rd row and the 3rd column, these row and column ordinal numbers are taken as coordinate points (such as (3, 3)), and then a second mapping curve is generated by connecting these points through a suitable mathematical method. Although the second mapping curve has the same slope due to the same horizontal and vertical coordinates, the length of the second mapping curve will vary because the row and column numbers of the screened elements in the matrix are different.
[0151] In this embodiment, the third threshold value is also a preset value, which is used to screen elements greater than the value from all non-diagonal elements in the semantic and position matching matrix. Its role is to focus on those elements that are not in the key corresponding position (non-diagonal line position) but still have a high matching degree. For example, the third threshold value is set to 0.6.
[0152] In this embodiment, the third element sequence is a sequence obtained by arranging all non-diagonal elements greater than the third threshold value in the semantic and position matching matrix in descending order. For example, there are 0.7, 0.63, 0.58, etc. in the non-diagonal elements, and the elements greater than the third threshold value 0.6 are sorted in descending order as [0.7, 0.63], which is the third element sequence.
[0153] In this embodiment, the third mapping curve is generated based on the row and column ordinal numbers of all elements in the third element sequence in the semantic and position matching matrix. Similarly, the curve is constructed according to the row and column position information of the elements in the third element sequence in the matrix. For example, the element 0.7 in the third element sequence is located in the 1st row and the 2nd column in the matrix, and 0.63 is located in the 2nd row and the 4th column in the matrix. These row and column ordinal numbers are taken as coordinate points (such as (1, 2), (2, 4), etc.), and a third mapping curve is generated by connecting these points through a suitable mathematical method.
[0154] In this embodiment, the first (second, third) standard mapping curve is a curve with specific characteristics that is preset as a reference standard for measuring the matching degree of the first (second, third) mapping curve. These standard mapping curves are determined according to the design goals of the system, the characteristics of the knowledge field, and a large amount of experimental or empirical data. For example, in a certain specific knowledge field, a curve representing the distribution characteristics of elements with a higher matching degree in an ideal matching state is determined as the first standard mapping curve through analysis of a large amount of sample data. It provides a benchmark for evaluating the closeness of the actual generated mapping curve to the ideal matching situation.
[0155] In this embodiment, the matching degree of the first mapping curve and the first standard mapping curve, the matching degree of the second mapping curve and the second standard mapping curve, and the matching degree of the third mapping curve and the third standard mapping curve are calculated by a specific algorithm (for example, a curve similarity measurement algorithm can be used, such as calculating the Euclidean distance between two curves, the cosine similarity, etc. to quantify the matching degree).
[0156] The above technology has the following beneficial effects: By screening and sorting the semantic and position matching matrix elements according to different thresholds to generate mapping curves, the matching matrix information can be deeply analyzed from multiple angles. The first mapping curve is generated by sorting elements greater than the first threshold, which can comprehensively grasp the distribution of elements with higher matching degree and understand the prominent matching situation between the user query and the knowledge item. The second mapping curve is generated by sorting diagonal elements greater than the second threshold, which focuses on the core position matching relationship and accurately evaluates the key corresponding position matching degree. The third mapping curve is generated by sorting non-diagonal elements greater than the third threshold, which excavates potential matching relationships across positions and enriches the matching analysis dimension. The matching degree between the generated mapping curve and the standard mapping curve is used as a multi-dimensional matching score, which provides a scientific and quantitative evaluation standard. By comparing the actual curve with the standard curve, the matching degree can be intuitively and accurately measured, making the score more persuasive and reliable. This method avoids the one-sidedness of single-dimensional analysis, fully reflects the complex semantic and position relationship, improves the knowledge matching accuracy and depth, provides accurate and detailed knowledge matching results for users, and optimizes the quality of knowledge retrieval service and user experience.
[0157] Embodiment 8: Based on embodiment 1, a comprehensive weight matching module is added, which refers to Figure 4 , comprising:
[0158] A semantic and position constraint degree determination submodule is configured to filter all synonymous entity combinations in the domain knowledge graph of the user query, and determine the semantic and position constraint degrees of all synonymous entity combinations based on the sorting values of all synonymous entity combinations in all reference knowledge items.
[0159] A weight assignment submodule is configured to assign weights to the multi-dimensional matching scores based on the semantic and position constraint degrees of all synonymous entity combinations to obtain multi-dimensional weights of the multi-dimensional matching scores.
[0160] A calibrated matching score generation submodule is configured to add the multi-dimensional matching scores according to the multi-dimensional weights to obtain the calibrated matching scores of the user query and the corresponding structured comprehensive character class knowledge item, and take the calibrated matching scores of all structured comprehensive character class knowledge items in the knowledge base as the matching results.
[0161] The first knowledge item screening submodule is configured to take all the structured comprehensive character class knowledge items in the knowledge base with the calibration matching scores greater than the first score threshold as all target knowledge items, and generate a semantic association graph of the user query and all the structured comprehensive character class knowledge items;
[0162] The second knowledge item screening submodule is configured to take all the structured comprehensive character class knowledge items in the knowledge base with the calibration matching scores greater than the second score threshold and not greater than the first score threshold as all expanded recommended knowledge items.
[0163] The personalized retrieval report generation submodule is configured to generate a personalized retrieval report based on all the target knowledge items and all the expanded recommended knowledge items.
[0164] In this embodiment, screening all the near-synonymous entity combinations in the domain knowledge graph of the user query means finding the sets of entities with similar semantics in the specific domain knowledge graph related to the user query. For example, when the user queries the medical field related content, “high blood pressure” and “blood pressure rise” may constitute a near-synonymous entity combination in the medical knowledge graph.
[0165] In this embodiment, the reference knowledge item refers to the knowledge item in the domain of the user query, which is used to reference and analyze the information related to the near-synonymous entity combination. These knowledge items contain at least one entity (i.e., sub-information body) in the near-synonymous entity combination.
[0166] In this embodiment, the ranking value of the near-synonymous entity combination in all the reference knowledge items refers to the order value of the entity (or sub-information body) in the near-synonymous entity combination in the reference knowledge items.
[0167] In this embodiment, based on the ranking values of all the near-synonymous entity combinations in all the reference knowledge items, the semantic and position constraint degrees of all the near-synonymous entity combinations are determined, that is, by analyzing the similarity of the ranking values of the two entities (or sub-information bodies) in the near-synonymous entity combination in the corresponding reference knowledge items, and combining a certain calculation method to determine their constraint degrees in the semantic and position aspects.
[0168] In this embodiment, the calibration matching score of the user query and the corresponding structured comprehensive character class knowledge item is obtained by weighting and adding the multi-dimensional matching scores based on the multi-dimensional weighting, which means multiplying the dimension values of the multi-dimensional matching scores by the corresponding weights and then adding them to obtain a comprehensive calibration matching score.
[0169] In this embodiment, the calibrated matching score is obtained by the above-mentioned weighted summation method, and is a numerical value for measuring the matching degree of the user query and the corresponding structured comprehensive character class knowledge item. It integrates multi-dimensional matching scores and weight information of each dimension, and is an important basis for judging the relevance of the knowledge item and the user query.
[0170] In this embodiment, the first score threshold is a pre-set numerical standard for screening knowledge items with high matching degrees with the user query, for example, the first score threshold is set to 0.8.
[0171] In this embodiment, the target knowledge item refers to a structured comprehensive character class knowledge item in the knowledge base with a calibrated matching score greater than the first score threshold. These knowledge items have a high matching degree with the user query and are considered to be the core knowledge content most consistent with the user's needs. The system will preferentially present such knowledge to the user.
[0172] In this embodiment, the semantic association graph of the user query and all structured comprehensive character class knowledge items is a graph that displays the semantic relationship between the user query and all structured comprehensive character class knowledge items in the knowledge base in a graphical manner. In this graph, nodes can represent key entities in the user query and related entities in the knowledge item, and edges represent the semantic relationship between these entities, such as similarity relationship, causal relationship, etc. For example, for the user query "Application of artificial intelligence in image recognition", the graph takes "artificial intelligence", "image recognition", "application case", etc. as nodes and displays their association with entities in each knowledge item in the knowledge base through edges.
[0173] In this embodiment, the second score threshold is also a pre-set numerical value, which is less than the first score threshold. Its function is to screen knowledge items that have some relevance with the user query but have a slightly lower matching degree than the target knowledge item. For example, the second score threshold is set to 0.6.
[0174] In this embodiment, the expanded recommended knowledge item refers to a structured comprehensive character class knowledge item in the knowledge base with a calibrated matching score greater than the second score threshold and not greater than the first score threshold.
[0175] In this embodiment, the personalized search report is generated based on all target knowledge items and all expanded recommended knowledge items, that is, the system organizes and presents the target knowledge items and the expanded recommended knowledge items to the user in a personalized manner. The report may include detailed content of the target knowledge items, brief introduction of the expanded recommended knowledge items, and the above-mentioned semantic association graph and other information.
[0176] The beneficial effects of the above technology are: the semantic and position constraint degree determination submodule determines the semantic and position constraint degree by screening the near-sense entity combination in the domain knowledge graph to which the user query belongs, deeply mines the internal relationship between knowledge, and provides accurate basis for weight assignment. The weight assignment submodule assigns weights to the multi-dimensional matching scores based on the semantic and position constraint degree, makes the weights scientific and reasonable, and highlights the key knowledge item matching score. The calibrated matching score generation submodule adds the multi-dimensional weights and multi-dimensional matching scores according to the weights to obtain the calibrated matching score, so as to greatly improve the matching accuracy. The first and second knowledge item screening submodules screen the target knowledge items and the expanded recommended knowledge items in different score thresholds, meet the user's core and expanded knowledge needs. The personalized search report generation submodule generates a personalized search report based on the screened items, optimizes the user search experience, and improves the system practicability and user satisfaction.
[0177] Embodiment 9: Based on embodiment 8, the semantic and position constraint degree determination submodule, referring to Figure 4 , comprises:
[0178] The semantic and position constraint matrix generation submodule is configured to screen all near-sense entity combinations in the domain knowledge graph to which the user query belongs.
[0179] The semantic and position constraint degree calculation submodule is configured to determine the semantic and position constraint degree of each near-sense entity combination based on the comprehensive similarity of the ordering value of each near-sense entity combination in all reference knowledge items and the near-sense degree.
[0180] In this embodiment, the comprehensive similarity of the ordering value of the near-sense entity combination in all reference knowledge items is the ratio of the average value of the ordering value of the two entities (sub-information bodies) in the near-sense entity combination in all reference knowledge items.
[0181] In this embodiment, the near-sense degree of the near-sense entity combination mainly measures the degree of similarity of the semantics of each entity in the near-sense entity combination. It is obtained by searching in a pre-prepared list containing the near-sense degree of all near-sense entity combinations.
[0182] In this embodiment, the semantic and position constraint degree of each near-sense entity combination is determined based on the comprehensive similarity of the ordering value of each near-sense entity combination in all reference knowledge items and the near-sense degree: that is, the product of the comprehensive similarity and the near-sense degree of each near-sense entity combination in all reference knowledge items is taken as the semantic and position constraint degree of each near-sense entity combination.
[0183] The beneficial effect of the above technology is that the semantic and position constraint degrees are determined by comprehensively considering the comprehensive similarity of the ordering values of each near-synonymous entity combination in all the reference knowledge entries and the near-synonym degree, which comprehensively and meticulously excavates the potential connection between entities in the knowledge entries. The comprehensive similarity reflects the similarity between different knowledge entries, and the near-synonym degree highlights the semantic relevance of the near-synonymous entity combination itself, and the combination of the two makes the determined semantic and position constraint degrees more accurately reflect the internal logic between knowledge. This not only helps to improve the understanding accuracy of the knowledge entries, but also provides more accurate weight basis for the subsequent weight assigning sub-module, thereby further optimizing the calculation of the calibration matching score and improving the accuracy and reliability of the knowledge matching result, ultimately providing the user with a knowledge retrieval result that better meets the user's needs and enhancing the overall performance of the knowledge retrieval system.
[0184] Embodiment 10: Based on Embodiment 8, the weight assigning sub-module, according to Figure 4 , comprises:
[0185] The near-synonymous entity combination statistical unit is configured to count the number of all near-synonymous entity combinations with a semantic and position constraint degree greater than a constraint degree threshold as a first number, and simultaneously count the number of all near-synonymous entity combinations with a semantic and position constraint degree not greater than the constraint degree threshold as a second number.
[0186] The multi-dimensional weight assigning unit is configured to obtain a multi-dimensional assigning weight of the multi-dimensional matching score based on the first number and the second number, comprising:
[0187] The weight of the matching degree of the first mapping curve and the first standard mapping curve is set to 0.5, the product of the ratio of the first number to the total number of all near-synonymous entity combinations and 0.5 is taken as the weight of the matching degree of the second mapping curve and the second standard mapping curve, and the product of the ratio of the second number to the total number of all near-synonymous entity combinations and 0.5 is taken as the weight of the matching degree of the third mapping curve and the third standard mapping curve.
[0188] In this embodiment, the constraint degree threshold is a pre-set standard value for classifying or screening the semantic and position constraint degrees of the near-synonymous entity combinations. For example, the constraint degree threshold is set to 0.7.
[0189] The beneficial effect of the above technology is that through such weight setting, the multi-dimensional matching score can more accurately reflect the matching degree of the user query and the structured comprehensive character class knowledge entries, effectively improving the accuracy of knowledge matching, thereby improving the quality of knowledge retrieval services, presenting the user with knowledge content that better meets their needs, optimizing the user experience, and making the knowledge retrieval system more practical and reliable.
[0190] It will be apparent to those skilled in the art that various modifications and variations can be made to the present application without departing from the spirit or scope of the application. Thus, it is intended that the present application cover modifications and variations of this application provided they come within the scope of the application and their equivalent technology.
Claims
1. A vectorized knowledge representation and knowledge base system matching large models, characterized in that, The method comprises the following steps: a structured analysis module is configured to perform structured analysis on all unstructured comprehensive character class knowledge items in the knowledge base to obtain all structured comprehensive character class knowledge items; a knowledge item vectorization module is configured to convert each structured comprehensive character class knowledge item into a high-dimensional vector based on a vectorization model to obtain a comprehensive character class knowledge item high-dimensional vector of each structured comprehensive character class knowledge item; a vector multidimensional matching module is configured to generate a user query vector and calculate a multidimensional matching score of the user query vector and the comprehensive character class knowledge item high-dimensional vector of each comprehensive character class knowledge item in the knowledge base; a comprehensive weight matching module is configured to assign weights to the multidimensional matching score based on the semantics and position constraints of all synonymous entity combinations in the domain knowledge graph to which the user query belongs, obtain multidimensional weight assignment of the multidimensional matching score, obtain a matching result based on the multidimensional matching score and the multidimensional weight assignment, and generate a personalized search report based on the matching result; The comprehensive weight matching module comprises: a weight assignment sub-module, which comprises: a synonymous entity combination statistical unit is configured to count the number of all synonymous entity combinations with a semantic and position constraint degree greater than a constraint degree threshold as a first number, and simultaneously count the number of all synonymous entity combinations with a semantic and position constraint degree not greater than the constraint degree threshold as a second number; a multidimensional weight assignment unit is configured to obtain multidimensional weight assignment of the multidimensional matching score based on the first number and the second number, comprising: setting the weight of the matching degree of the first mapping curve and the first standard mapping curve to 0.5, and setting the product of the ratio of the first number to the total number of all synonymous entity combinations and 0.5 as the weight of the matching degree of the second mapping curve and the second standard mapping curve, and setting the product of the ratio of the second number to the total number of all synonymous entity combinations and 0.5 as the weight of the matching degree of the third mapping curve and the third standard mapping curve; The determination method of the first mapping curve, the second mapping curve and the third mapping curve comprises: based on the semantic and position matching score of each sub-information body of the user query and each sub-information body in each structured comprehensive character class knowledge item, a semantic and position matching matrix of the user query and each structured comprehensive character class knowledge item is constructed; all elements greater than a first threshold, all diagonal elements greater than a second threshold, and all non-diagonal elements greater than a third threshold in the semantic and position matching matrix of the user query and each structured comprehensive character class knowledge item are sorted from large to small to obtain a first element sequence, a second element sequence, and a third element sequence, and based on the row and column ordinal numbers of all elements in the semantic and position matching matrix in the first element sequence, the second element sequence, and the third element sequence, a first mapping curve, a second mapping curve, and a third mapping curve are respectively generated.
2. The vectorized knowledge representation and large model matching knowledge base system of claim 1, wherein, The knowledge item vectorization module comprises: a knowledge item encoding sub-module is configured to perform absolute position encoding and relative position association encoding on all sub-information bodies in each structured comprehensive character class knowledge item to generate a position and semantic perception vector of each sub-information body; The feature alignment submodule is configured to align and map the position and semantic awareness vectors of the sub-information bodies of all character classes in each structured comprehensive character class knowledge item to the same feature space, and obtain a comprehensive character class knowledge item high-dimensional vector of each structured comprehensive character class knowledge item.
3. The vectorized knowledge representation and large model matched knowledge base system of claim 2, wherein, The knowledge item encoding submodule comprises: a basic encoding unit configured to divide each sub-information body of all character classes in each structured comprehensive character class knowledge item, encode each sub-information body based on the semantic content of each sub-information body, and obtain a basic encoding of each sub-information body; an absolute position encoding unit configured to generate an absolute position encoding of each sub-information body based on the different character class position and the same character class position of each sub-information body; a relative position encoding unit configured to generate a relative position association encoding of each sub-information body based on the same character class relative association position and the different character class relative association position of each sub-information body; a position and semantic awareness vectorization unit configured to generate a position and semantic awareness vector of each sub-information body based on the basic encoding, the absolute position encoding, the relative position encoding of each sub-information body, and a vectorization model.
4. The vectorized knowledge representation and large model matching knowledge base system of claim 3, wherein, The absolute position encoding unit comprises: a different character class positioning subunit configured to determine the sorting position of each sub-information body in the corresponding structured comprehensive character class knowledge item, and generate a different character class position of each sub-information body based on the sorting position of each sub-information body in the corresponding structured comprehensive character class knowledge item; an information body classification subunit configured to sequentially reserve and classify the same character class sub-information bodies in each sub-information body of all character classes in each structured comprehensive character class knowledge item, and obtain a same character class sub-information body set of each structured comprehensive character class knowledge item; a same character class positioning subunit configured to determine the sorting position of each sub-information body in the corresponding same character class sub-information body set, and generate a same character class position of each sub-information body based on the corresponding same character class sub-information body set and the sorting position of each sub-information body in the corresponding same character class sub-information body set; an absolute position encoding subunit configured to generate an absolute position encoding of each sub-information body based on the different character class position and the same character class position of each sub-information body and a preset absolute position encoding mode.
5. The vectorized knowledge representation and large model matching knowledge base system of claim 3, wherein, The relative position encoding unit comprises: a same character class relative association positioning subunit configured to take the sorting positions of all synonymous sub-information bodies of each sub-information body of each structured comprehensive character class knowledge item in the corresponding same character class sub-information body set as the same character class relative association position of each sub-information body; a different character class relative association positioning subunit configured to take the sorting positions of all synonymous sub-information bodies of each sub-information body of each structured comprehensive character class knowledge item in the corresponding sub-information body of all character classes as the different character class relative association position of each sub-information body; a relative position encoding subunit configured to generate a relative position association encoding of each sub-information body based on the same character class relative association position and the different character class relative association position of each sub-information body and a preset relative position association encoding mode.
6. The vectorized knowledge representation and large model matching knowledge base system of claim 1, wherein, The vector multidimensional matching module comprises: The information body vectorization submodule is configured to determine semantic and location-aware vectors of each sub-information body of the user query based on the user query vector, and determine semantic and location-aware vectors of each sub-information body in each structured comprehensive character class knowledge item based on the high-dimensional vector of each comprehensive character class knowledge item in the knowledge base. The semantic and location matching submodule is configured to calculate semantic and location matching scores of each sub-information body of the user query and each sub-information body in each structured comprehensive character class knowledge item based on similarities of elements at all same positions in the semantic and location-aware vectors of each sub-information body of the user query and the semantic and location-aware vectors of each sub-information body in each structured comprehensive character class knowledge item. The matching score matrixing submodule is configured to construct a semantic and location matching matrix of the user query and each structured comprehensive character class knowledge item based on the semantic and location matching scores of each sub-information body of the user query and each sub-information body in each structured comprehensive character class knowledge item. The multidimensional matching submodule is configured to construct mapping curves based on all diagonal elements and all non-diagonal elements of the semantic and location matching matrix of the user query and each structured comprehensive character class knowledge item, and calculate multidimensional matching scores of the user query vector and the high-dimensional vector of each comprehensive character class knowledge item in the knowledge base.
7. The vectorized knowledge representation and large model matching knowledge base system of claim 6, wherein, The multidimensional matching submodule comprises: The first mapping curve generation unit is configured to sort, in descending order, all elements greater than a first threshold value in all elements of the semantic and location matching matrix of the user query and each structured comprehensive character class knowledge item, to obtain a first element sequence, and generate a first mapping curve based on row and column serial numbers of all elements in the first element sequence in the semantic and location matching matrix. The second mapping curve generation unit is configured to sort, in descending order, all elements greater than a second threshold value in all diagonal elements of the semantic and location matching matrix of the user query and each structured comprehensive character class knowledge item, to obtain a second element sequence, and generate a second mapping curve based on row and column serial numbers of all elements in the second element sequence in the semantic and location matching matrix. The third mapping curve generation unit is configured to sort, in descending order, all elements greater than a third threshold value in all non-diagonal elements of the semantic and location matching matrix of the user query and each structured comprehensive character class knowledge item, to obtain a third element sequence, and generate a third mapping curve based on row and column serial numbers of all elements in the third element sequence in the semantic and location matching matrix. The multidimensional matching degree calculation unit is configured to take matching degrees of the first mapping curve and a first standard mapping curve, matching degrees of the second mapping curve and a second standard mapping curve, and matching degrees of the third mapping curve and a third standard mapping curve as the multidimensional matching scores of the user query vector and the high-dimensional vector of each comprehensive character class knowledge item in the knowledge base.
8. The vectorized knowledge representation and large model matching knowledge base system of claim 1, wherein, The comprehensive weight matching module comprises: The semantic and position constraint degree determination submodule is configured to filter all synonym entity combinations in the domain knowledge graph of the user query, and determine the semantic and position constraint degrees of all synonym entity combinations based on the ranking values of all synonym entity combinations in all reference knowledge entries; The weight assignment submodule is configured to assign weights to the multi-dimensional matching scores based on the semantic and position constraint degrees of all synonym entity combinations to obtain multi-dimensional weights of the multi-dimensional matching scores; The calibrated matching score generation submodule is configured to sum the multi-dimensional matching scores based on the multi-dimensional weights to obtain the calibrated matching scores of the user query and the corresponding structured comprehensive character class knowledge entries, and take the calibrated matching scores of all structured comprehensive character class knowledge entries in the knowledge base as matching results; The first knowledge entry filtering submodule is configured to take all structured comprehensive character class knowledge entries in the knowledge base with calibrated matching scores greater than a first score threshold as all target knowledge entries, and generate a semantic association graph of the user query and all structured comprehensive character class knowledge entries; The second knowledge entry filtering submodule is configured to take all structured comprehensive character class knowledge entries in the knowledge base with calibrated matching scores greater than a second score threshold and not greater than the first score threshold as all expanded recommended knowledge entries; The personalized search report generation submodule is configured to generate a personalized search report based on all target knowledge entries and all expanded recommended knowledge entries.
9. The vectorized knowledge representation and large model matching knowledge base system of claim 8, wherein, The semantic and position constraint degree determination submodule includes: The semantic and position constraint matrix generation submodule is configured to filter all synonym entity combinations in the domain knowledge graph of the user query; The semantic and position constraint degree calculation submodule is configured to determine the semantic and position constraint degrees of each synonym entity combination based on the comprehensive similarity and synonym degree of the ranking values of each synonym entity combination in all reference knowledge entries.
Citation Information
Patent Citations
Semantic representation method for patent text vectors
CN104199809A
Text-pedestrian retrieval method based on bounding box extraction and semantic consistency constraint
CN116842212A