Semantic retrieval methods and devices

CN122088508BActive Publication Date: 2026-09-01BEIJING ZHONGZHI SMART TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202512033560.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-09-01
Estimated Expiration
2045-12-30

AI Technical Summary

Technical Problem

虽然一定程度上提升了语义层面的精度,但由于大模型输入窗口有限、计算资源消耗大,在面对大规模语料时依然存在漏检与效率低下的问题

Benefits of technology

[0025]首先,本发明实施例根据历史检索内容与历史关键实体信息之间的样本数据预先训练生成原子条件识别模型,在进行实时语义检索时,将用户输入检索内容输入至该训练好的原子条件识别小模型,提取出关键实体信息,生成原子条件信息表,每一关键实体信息对应一原子条件。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122088508B_ABST
    Figure CN122088508B_ABST
Patent Text Reader

Abstract

This invention discloses a semantic retrieval method and apparatus. The method includes: inputting search content into an atomic condition recognition model to generate an atomic condition information table; inputting the semantic information of the search content into an external general model to obtain expression pseudocode; inputting atomic conditions, search content, and expression pseudocode into the external general model to obtain atomic conditions with classification types, and extracting and combining them according to semantic information to obtain the relationship between atomic conditions; extracting search elements from the semantic information of the search content; obtaining classification limiting conditions according to the expanded entity element classification, and integrating them into the atomic condition information table to obtain updated relationships between atomic conditions; establishing triplet, quadruple, or multi-tuple models to form a vector model according to different application requirements and expanded search entity elements; and fusing the search results according to the vector model and the relationship between the updated atomic conditions. This invention can perform semantic retrieval efficiently and accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data technology, and in particular to a semantic retrieval method and apparatus. Background Technology

[0002] This section is intended to provide background or context for the embodiments of the invention set forth in the claims. The description herein is not an admission that it is prior art simply because it is included in this section.

[0003] There are two existing methods for semantic retrieval:

[0004] Statistical vector comparison methods, typically represented by TF-IDF and cosine similarity, offer advantages in terms of simplicity and speed. However, their fundamental drawback lies in neglecting word order and grammatical structure, leading to documents that are lexically similar but semantically unrelated being easily classified as similar.

[0005] Semantic enhancement methods based on large models typically involve first filtering candidate documents using vector alignment, and then inputting them into a large model for re-ranking. While this improves semantic accuracy to some extent, the large model still suffers from problems of missed detections and low efficiency when dealing with large-scale corpora due to its limited input window and high computational resource consumption.

[0006] Therefore, existing methods are either not accurate enough or cannot be scaled up for large-scale applications, resulting in low efficiency. Summary of the Invention

[0007] This invention provides a semantic retrieval method for efficient and accurate semantic retrieval, the method comprising:

[0008] The user inputs the search content into the atomic condition recognition model, extracts key entity information, and generates an atomic condition information table, with each key entity information corresponding to an atomic condition; the atomic condition recognition model is pre-trained based on sample data between historical search content and historical key entity information.

[0009] The semantic information of the user's search query is input into an external general model to obtain the pseudocode of the expression corresponding to the user's search query. The atomic conditions in the atomic condition information table, the user's search query, and the pseudocode of the expression are arranged together as prompts and input into the external general model for classification to obtain atomic conditions with classification types. The external general model is pre-trained and generated based on the sample data of the relationship between the search query and the pseudocode of the expression, and the sample data of the relationship between the prompts and the atomic conditions with classification types.

[0010] The relationships between atomic conditions are obtained by extracting and combining the semantic information of the user-input search content based on the atomic conditions with classification types.

[0011] Extract search elements from the semantic information of the search content entered by the user, expand the search elements into entities; classify the search entity elements according to the expanded search entity elements, obtain the search entity element classification constraints, integrate them into the atomic condition information table to obtain atomic conditions with classification types, and re-extract and combine them to obtain the updated relationships between atomic conditions.

[0012] Based on different application needs, the expanded retrieval entity elements are used to establish a vector model consisting of triples, quadruples, or multi-tuples. Triples are used to record the relationship between the subject and the object, quadruples are used to record the relationship between the subject, subject attributes, and the object, and multi-tuples are used to record the relationship between the subject, subject attributes, the object, and the object attributes, as well as the conditions required to generate the relationship.

[0013] Based on the relationship between the vector model and the updated atomic conditions, the semantic retrieval results are obtained through fusion retrieval.

[0014] This invention also provides a semantic retrieval device for efficient and accurate semantic retrieval, the device comprising:

[0015] The atomic condition information table generation unit is used to input the user's search content into the atomic condition recognition model, extract key entity information, and generate an atomic condition information table, where each key entity information corresponds to an atomic condition; the atomic condition recognition model is pre-trained and generated based on sample data between historical search content and historical key entity information.

[0016] The atomic classification unit is used to input the semantic information of the user-input search content into an external general model to obtain the expression pseudocode corresponding to the user-input search content; the atomic conditions in the atomic condition information table, the user-input search content, and the expression pseudocode are arranged together as prompts and input into the external general model for classification to obtain atomic conditions with classification types; the external general model is pre-trained and generated based on the sample data of the relationship between the search content and the expression pseudocode, and the sample data of the relationship between the prompts and the atomic conditions with classification types.

[0017] Atomic assembly unit is used to extract and assemble atomic conditions with classification types according to the semantic information of the user input search content to obtain the relationship between atomic conditions;

[0018] The element-limited fusion unit is used to extract search elements from the semantic information of the search content input by the user, expand the search elements into entities, obtain the search entity element classification limit conditions according to the classification of the expanded search entity elements, and merge them into the atomic condition information table to obtain atomic conditions with classification type again, so as to re-extract and combine to obtain the relationship between updated atomic conditions.

[0019] The vector model building unit is used to build a vector model from the expanded retrieved entity elements according to different application requirements, using triplet, quadruple, or multi-tuple models. Triplets are used to record the relationship between the subject and the object, quadruplets are used to record the relationship between the subject, subject attributes, and object, and multi-tuples are used to record the relationship between the subject, subject attributes, object, and object attributes, as well as the conditions required to generate the relationship.

[0020] The fusion retrieval unit is used to obtain semantic retrieval results by fusing retrieval based on the relationship between the vector model and the updated atomic conditions.

[0021] This invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described semantic retrieval method.

[0022] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the semantic retrieval method described above.

[0023] This invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the semantic retrieval method described above.

[0024] The beneficial technical effects of the semantic retrieval scheme provided in this embodiment of the invention are:

[0025] First, in this embodiment of the invention, an atomic condition recognition model is pre-trained based on sample data between historical search content and historical key entity information. When performing real-time semantic retrieval, the user input search content is input into the pre-trained atomic condition recognition model to extract key entity information and generate an atomic condition information table, with each key entity information corresponding to an atomic condition.

[0026] Secondly, this embodiment of the invention also utilizes an external general-purpose large model pre-trained based on sample data of the relationship between the search content and the expression pseudocode, and sample data of the relationship between the prompt content and the atomic conditions with classification types. During real-time semantic retrieval, the semantic information of the user-input search content is input into the external general-purpose large model to obtain the expression pseudocode corresponding to the user-input search content; the atomic conditions in the atomic condition information table, the user-input search content, and the expression pseudocode are arranged together as prompt content and input into the external general-purpose large model for classification to obtain atomic conditions with classification types.

[0027] Based on the above, this embodiment of the invention proposes a semantic retrieval scheme that combines small models and large models.

[0028] Furthermore, in this embodiment of the invention, atomic conditions with classification types are extracted and combined according to the semantic information of the user-input search content to obtain the relationship between atomic conditions; search elements are extracted from the semantic information of the user-input search content, and entity expansion is performed on the search elements; based on the classification of the expanded search entity elements, search entity element classification constraints are obtained, which are then merged into the atomic condition information table to obtain atomic conditions with classification types again, so as to re-extract and combine to obtain the updated relationship between atomic conditions. This realizes the fusion of search entity element classification constraints and atomic condition information table, improves the accuracy of the relationship between atomic conditions, and thus improves the accuracy of subsequent semantic retrieval.

[0029] Furthermore, according to different application requirements, the extended retrieval entity elements in this embodiment of the invention establish a vector model composed of triples, quadruples, or multi-tuples. Triples are used to record the relationship between the subject and the object, quadruples are used to record the relationship between the subject, subject attributes, and the object, and multi-tuples are used to record the relationship between the subject, subject attributes, the object, and object attributes, as well as the conditions required to generate the relationship. The vector model composed of triples, quadruples, or multi-tuples can improve the accuracy of subsequent semantic retrieval.

[0030] Finally, based on the relationship between the vector model and the updated atomic conditions, the present invention re-integrates the retrieval to obtain semantic retrieval results, thus realizing a semantic retrieval optimization method based on thought chain reasoning.

[0031] In summary, this invention proposes a semantic retrieval optimization method that combines small and large models and is based on thought chain reasoning, which can achieve efficient and accurate semantic retrieval in large-scale data environments. Attached Figure Description

[0032] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:

[0033] Figure 1 This is a flowchart illustrating the semantic retrieval method in an embodiment of the present invention;

[0034] Figure 2 This is a schematic diagram illustrating the principle of semantic retrieval in an embodiment of the present invention;

[0035] Figure 3 This is a schematic diagram of the process of integrating the classification constraints of the retrieved entity elements into the atomic condition information table in an embodiment of the present invention;

[0036] Figure 4 This is a schematic diagram of the semantic retrieval device in an embodiment of the present invention;

[0037] Figure 5 This is a schematic diagram of a computer device structure according to an embodiment of the present invention. Detailed Implementation

[0038] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.

[0039] The acquisition, storage, use, and processing of data in this application comply with relevant laws and regulations.

[0040] This invention pertains to the field of information retrieval and text processing technology, and relates to a semantic retrieval scheme. This scheme is a natural language retrieval optimization scheme based on a large model and thought chain, designed to improve retrieval accuracy in scenarios where users perform semantic searches using natural language on massive amounts of documents, enabling users to quickly locate target documents without reading a large number of documents. The semantic retrieval scheme is described in detail below.

[0041] Figure 1 This is a flowchart illustrating the semantic retrieval method in an embodiment of the present invention, such as... Figure 1 As shown, the method includes:

[0042] Step 101: Input the user's search content into the atomic condition recognition model, extract key entity information, and generate an atomic condition information table. Each key entity information corresponds to an atomic condition. The atomic condition recognition model is pre-trained based on sample data between historical search content and historical key entity information.

[0043] Step 102: Input the semantic information of the user-input search content into an external general-purpose large model to obtain the expression pseudocode corresponding to the user-input search content; Compile the atomic conditions in the atomic condition information table, the user-input search content, and the expression pseudocode together into prompt content and input it into the external general-purpose large model for classification to obtain atomic conditions with classification types; The external general-purpose large model is pre-trained and generated based on the sample data of the relationship between the search content and the expression pseudocode, and the sample data of the relationship between the prompt content and the atomic conditions with classification types.

[0044] Step 103: Extract and combine atomic conditions with classification types according to the semantic information of the user-input search content to obtain the relationship between atomic conditions;

[0045] Step 104: Extract search elements from the semantic information of the user-input search content, expand the search elements into entities; based on the expanded search entity element classification, obtain the search entity element classification constraints, integrate them into the atomic condition information table to obtain atomic conditions with classification types, and re-extract and combine them to obtain the updated relationships between atomic conditions.

[0046] Step 105: Based on different application requirements, the expanded retrieval entity elements are used to establish a vector model consisting of triples, quadruples, or multi-tuples. Triples are used to record the relationship between the subject and the object, quadruples are used to record the relationship between the subject, subject attributes, and the object, and multi-tuples are used to record the relationship between the subject, subject attributes, the object, and the object attributes, as well as the conditions required to generate the relationship.

[0047] Step 106: Based on the relationship between the vector model and the updated atomic conditions, the semantic retrieval results are obtained through fusion retrieval.

[0048] The beneficial technical effects of the semantic retrieval method provided in this embodiment of the invention are:

[0049] First, in this embodiment of the invention, an atomic condition recognition model is pre-trained based on sample data between historical search content and historical key entity information. When performing real-time semantic retrieval, the user input search content is input into the pre-trained atomic condition recognition model to extract key entity information and generate an atomic condition information table, with each key entity information corresponding to an atomic condition.

[0050] Secondly, this embodiment of the invention also utilizes an external general-purpose large model pre-trained based on sample data of the relationship between the search content and the expression pseudocode, and sample data of the relationship between the prompt content and the atomic conditions with classification types. During real-time semantic retrieval, the semantic information of the user-input search content is input into the external general-purpose large model to obtain the expression pseudocode corresponding to the user-input search content; the atomic conditions in the atomic condition information table, the user-input search content, and the expression pseudocode are arranged together as prompt content and input into the external general-purpose large model for classification to obtain atomic conditions with classification types.

[0051] Based on the above, this embodiment of the invention proposes a semantic retrieval scheme that combines small models and large models.

[0052] Furthermore, in this embodiment of the invention, atomic conditions with classification types are extracted and combined according to the semantic information of the user-input search content to obtain the relationship between atomic conditions; search elements are extracted from the semantic information of the user-input search content, and entity expansion is performed on the search elements; based on the classification of the expanded search entity elements, search entity element classification constraints are obtained, which are then merged into the atomic condition information table to obtain atomic conditions with classification types again, so as to re-extract and combine to obtain the updated relationship between atomic conditions. This realizes the fusion of search entity element classification constraints and atomic condition information table, improves the accuracy of the relationship between atomic conditions, and thus improves the accuracy of subsequent semantic retrieval.

[0053] Furthermore, according to different application requirements, the extended retrieval entity elements in this embodiment of the invention establish a vector model composed of triples, quadruples, or multi-tuples. Triples are used to record the relationship between the subject and the object, quadruples are used to record the relationship between the subject, subject attributes, and the object, and multi-tuples are used to record the relationship between the subject, subject attributes, the object, and object attributes, as well as the conditions required to generate the relationship. The vector model composed of triples, quadruples, or multi-tuples can improve the accuracy of subsequent semantic retrieval.

[0054] Finally, based on the relationship between the vector model and the updated atomic conditions, the present invention re-integrates the retrieval to obtain semantic retrieval results, thus realizing a semantic retrieval optimization method based on thought chain reasoning.

[0055] In summary, this invention proposes a semantic retrieval optimization method that combines small and large models and is based on thought chain reasoning, enabling efficient and accurate semantic retrieval in large-scale data environments. The following section will further elaborate on this method. Figures 2 to 4 This semantic retrieval method will be described in detail.

[0056] First, the embodiments of the present invention perform condition analysis, which can be achieved through... Figure 2 The "Condition Analysis Module" in the system is used to achieve this: a two-layer analytical model is constructed that combines a dedicated small model with a general large model. By identifying, classifying and combining atomic conditions in the user input information, the system can accurately interpret the user's intent.

[0057] For example, a user inputs their search query (content): "Search for the 10 most similar patents in country A for soy milk maker pressure relief devices that are most similar to CN102397011A".

[0058] In step 101 above, as Figure 2The "Atomic Condition Recognition" shown refers to the extraction of key entity information from user input using a dedicated recognition model (atomic condition recognition model, small model) trained on historical user question-and-answer pairs. This generates an atomic condition information table, which serves as the basis for subsequent retrieval constraints. This key entity information includes, but is not limited to, technical topic, technical field, data scope, time range, patent type, legal status, and patent number. If the user input does not explicitly specify required information, default values ​​will be used. If the user does not specify a data scope, a search will be performed across the entire database by default.

[0059] In step 101 above, the atomic condition information table identified by the atomic condition recognition model is shown in Table 1 below:

[0060] Table 1: Atomic Condition Information Table

[0061] CN102397011A Publication Number (PN) Most similar to CN102397011A 10 (Results set) 10 items Soy milk maker Title (TI) Soy milk maker pressure relief device pressure relief device Title (TI) Soy milk maker pressure relief device Country A Country and Region (CA) Country A patent

[0062] In step 102 above, the semantic information of the user-input search content is input into an external general-purpose model to obtain the pseudocode of the expression corresponding to the user-input search content; the atomic conditions in the atomic condition information table, the user-input search content, and the pseudocode of the expression are arranged together as prompts and input into the external general-purpose model for classification to obtain atomic conditions with classification types; the external general-purpose model is pre-trained and generated based on sample data of the relationship between the search content and the pseudocode of the expression, and sample data of the relationship between the prompts and the atomic conditions with classification types. Following the example of step 101, the classification table of atomic conditions with classification types, that is, the atomic condition classification table finally generated by the model, is shown in Table 2 below:

[0063] CN102397011A Publication Number (PN) Most similar to CN102397011A Target constraints, semantics PN=similar_to(CN102397011A) 10 (Result set RS) 10 items Target constraints, Boolean RS=TOP(10) Soy milk maker Title (TI) Soy milk maker pressure relief device Content characteristics conditions, Boolean TI = contains (soy milk maker) pressure relief device Title (TI) Soy milk maker pressure relief device Content characteristics conditions, Boolean TI = contains (pressure relief device) Country A Country and Region (CA) Invention patent of country A Metadata filtering conditions, Boolean CA=(Country A)

[0064] Table 2: Classification Table of Atomic Conditions

[0065] In specific implementation, in step 102 above, such as Figure 2As shown in: "Atomic Condition Classification": constructing atomic condition classification standards and expression pseudocode rules, and accurately classifying and standardizing the expression of atomic conditions with the help of a large model. To express the user's retrieval intention more accurately, each condition in the atomic condition information table generated in the previous step is classified. In order to improve the classification accuracy and avoid identification errors, the atomic conditions, user input, and expression pseudocode rules are compiled together as prompt content and sent to an external general large model for classification. The model classifies atomic conditions into three categories: target restriction conditions (core objects and intentions of retrieval), content feature conditions (describing the technical content of patents per se), and metadata filtering conditions (screening and limiting results). That is, in one embodiment, the atomic conditions in the atomic condition information table, the content input for retrieval by the user, and the expression pseudocode rules are compiled together as prompt content and sent to an external general large model for classification, so as to obtain atomic conditions with classification types, including the following three types of atomic conditions: target restriction conditions, content feature conditions, and metadata filtering conditions. Meanwhile, by virtue of the understanding ability of the large model, the semantic information input by the user is converted into standardized expression pseudocode, which facilitates subsequent program processing. For ambiguous conditions that cannot be classified, they can be classified by default into similarity retrieval under target restriction conditions.

[0066] In the above step 103, as Figure 2 shown in "Atomic Condition Combination": extracting the relationships between atomic conditions based on semantic information, and mapping the relationships to corresponding association logics.

[0067] In specific implementation, the Boolean atomic conditions among the plurality of atomic conditions parsed in the above steps are compiled together with user input into prompt content and sent to an external large model, which performs relationship extraction according to the semantics of the user input and carries out corresponding logical combination for subsequent use. For example, "soybean milk maker" and "pressure reliever" should be combined into an AND relationship rather than an OR relationship.

[0068] In specific implementation, in the above step 103, for atomic condition combination, generally, Boolean conditions of different fields are connected by an AND operator, and multiple conditions of the same field are connected by an OR operator, and the combined Boolean expression is sent to the large model for verification and correction. The generated Boolean expression is as follows: CA=Country A AND (TI=soybean milk maker OR pressure reliever) AND PT=invention AND RS=TOP(10).

[0069] Secondly, as Figure 2 shown, the element analysis in the embodiment of the present invention can be implemented through Figure 2 the "element analysis module" in

[0070] In the above step 104, as Figure 2The “Retrieval Element Extraction” shown here extracts retrieval element words from semantic conditions.

[0071] The original text to be examined is obtained from semantic conditions, and keywords related to the technical solution are extracted using a large model.

[0072] Different weights are assigned based on the relevance of the technical solutions.

[0073] In step 104 above, the element words are expanded, such as... Figure 2 The “Search Element Entity Expansion” shown uses a local professional thesaurus and a large model to expand the search keywords.

[0074] Based on a local thesaurus, it enables synonym expansion (e.g., "computer" -> "computer"), near-synonym expansion (e.g., "corrosion" -> "erosion", "rot"), and hypernym and hyponym expansion (e.g., "atmospheric corrosion field" -> "atmospheric corrosion monitoring", "atmospheric corrosion sensor"), avoiding missed detections due to differences in expression.

[0075] The model is expanded with synonyms, near-synonyms, and translated terms to avoid missed detections due to expression and translation issues. The local keyword expansion list is merged and deduplicated with the large model expansion list to obtain the final keyword list. This list is then used to replace the corresponding keywords in the atomic conditions. For example, if the user inputs the keyword "pressure relief device" and the atomic condition is TI=pressure relief device, after expansion and replacement, it becomes TI=(pressure relief device or pressure relief apparatus or pressure relief valve or exhaust device). That is, "pressure relief device" in Table 1 above is replaced with: pressure relief device, pressure relief apparatus, pressure relief valve, exhaust device.

[0076] In step 104 above, the element word classification is limited, such as... Figure 2 The "Search Element Classification Limitation" shown below:

[0077] If the atomic condition information contains similar patent limitation conditions, the classification number of the similar patent limitation conditions is obtained through the query tool in the tool module and used as the limitation condition for retrieval, and then merged into the atomic condition information table.

[0078] If the atomic condition information contains only keywords, then the keyword vector is calculated and compared with the pre-vectorized classification table using an external classification query tool. The N categories with the highest vector similarity are returned as limiting conditions and merged into the atomic condition information table.

[0079] If the atomic condition information does not contain either similar patent limitation conditions or keywords, then the classification limitation conditions are ignored.

[0080] In one embodiment, in step 104 above, as follows: Figure 3As shown, based on the expanded retrieval entity element classification, retrieval entity element classification constraints are obtained, which are then integrated into the atomic condition information table to re-obtain atomic conditions with classification types. The relationships between these atomic conditions, which are then re-extracted and combined to obtain updated atomic conditions, can include:

[0081] Step 201: If the atomic condition information contains similar patent limitation conditions, obtain the classification number of the similar patent limitation conditions, use it as the classification limitation condition for the search entity element, and merge it into the atomic condition information table;

[0082] Step 202: If the atomic condition information contains only keywords, then by calculating the keyword vector, comparing it with the pre-vectorized classification table through an external classification query tool, and returning the N classifications with the highest vector similarity as the classification constraints for the retrieved entity elements, and merging them into the atomic condition information table;

[0083] Step 203: If the atomic condition information does not contain either similar patent limitation conditions or keywords, then ignore the classification limitation conditions.

[0084] In step 104 above, based on the semantic condition PN=similar_to(CN102397011A), the full text of the patent is obtained using an external tool. The result set is then limited according to the patent's classification number. To prevent missed detections, wildcards can be used to limit the classification number, generating the following expression: IPC=(A47J31% OR A23C11%).

[0085] Furthermore, the vector model is constructed in this embodiment of the invention through... Figure 2 This can be achieved using the "Vector Model Building Module" in the documentation:

[0086] In step 105 above, such as Figure 2 The “entity tuple construction” shown:

[0087] To address the logical relationships between entity keywords, an entity-relationship tuple model is established. Depending on different application requirements, triple, quadruple, or multi-tuple models, as well as weighted tuple models, can be created.

[0088] Triads record the relationship between a subject and an object, represented as (subject, object, relation), such as (soy milk maker, pressure relief device, contain). When the subject, object, and relation have synonyms, near-synonyms, hyponyms, or translated terms, a list is used, such as ([main frame, main frame], [sub-frame, sub-frame], [has, contains]).

[0089] A quadruple is used to record entities, emphasizing the attributes of the subject and the relationship between the subject and the object. It is represented as (subject, attribute, object, relationship), such as (soy milk maker, [aluminum, vertical], pressure relief device, [detachable], [has, contains]).

[0090] In addition to recording the subject, the subject's attributes, and the relationship between the subject and the object, tuples can also add the object's attributes, the conditions required to generate the relationship, etc., to generate higher-dimensional tuples.

[0091] In practice, the more element types a tuple contains, the more information it contains, and the more accurate the search results will be. However, the more complex it is to extract this information from the text and use it for retrieval. In practical applications, we can take into account both information extraction efficiency and search result accuracy, and flexibly select and compare suitable tuple types.

[0092] In step 105 above, the entity relationship weights are adjusted, such as... Figure 2 The “tuple weight adjustment” shown:

[0093] Supports relationship weight adjustment. By default, the weights of multiple attributes of an entity and multiple relationships between entities are the same. However, the interface provides users with the opportunity to adjust the weights and set different weights for different relationships, such as: (soy milk maker, pressure relief device, [aluminum], [removable (0.6), washable (0.4)], [includes]).

[0094] Tuple weights: Different parts of a long text usually have different importance, and tuples extracted from each part can also be assigned different weights.

[0095] As can be seen from the above, in one embodiment, the semantic retrieval method may further include: assigning different weights to triples, quadruples or multi-tuples, or assigning different weights to different elements in triples, quadruples or multi-tuples.

[0096] In step 105 above, to reduce the computational load, only the claims section of the patent document is used to extract the search elements. In practical applications, the elements can also be extracted from the specification, and the extracted entity-type search element terms can be expanded with synonyms, near-synonyms, etc. For example, pressure relief device can be expanded to exhaust device, pressure relief valve, safety valve, etc. The extracted entity relation tuples are shown in Tables 3(1) and 3(2) below:

[0097] Table 3(1): Entity Relationship Tuples

[0098]

[0099] Table 3(2): Entity Relationship Tuples

[0100]

[0101] In step 105 above, in a further preferred embodiment, such as Figure 2 The “Vector Computation Model Construction” shown:

[0102] (1) Decompose the tuple into its decomposition as follows: , where w i This refers to the weight of the i-th element. For example, w1 is the weight of the first element, w2 is the weight of the second element, and so on. i It represents the value of the i-th element, for example, e1 is the value of the first element, e2 is the value of the second element, and so on. The default weight is 1, and v represents a vector array of tuples.

[0103] (2) Calculate the vector of each element after tuple decomposition using a text embedding model (e.g., m3e, LM-L12, etc.). , where v i It is a vector with the i-th element.

[0104] (3) Combine the single vectors of the elements into a one-dimensional vector of the tuple. .

[0105] (4) Further aggregate the tuple vectors into text vectors, which can be calculated using the weighted average method.

[0106]

[0107] Here, vectors_array represents a vector array of tuples, and weights_array represents the weights of the tuples, which default to 1.

[0108] In a further preferred embodiment, the vector calculation method set in the technical solution is used to concatenate the above tuples after expansion with synonyms, near-synonyms, etc., to form an array to calculate the vector VectorToSearch, and thus obtain the vector retrieval expression R=VectorToSearch.

[0109] Furthermore, to improve the accuracy and flexibility of retrieval, tuples with the same weight can be aggregated into a text vector, and finally, multiple vector sets can be used to represent the content features of the text.

[0110] In step 105 above, in a further preferred embodiment, the vector distance is calculated as follows:

[0111] For single-vector comparison, the most commonly used text vector similarity metric is cosine similarity. In practical applications, if the text is represented by multiple vectors, different vectors can be used for comparison based on different input requirements. For example, patent text can be stored as vectors for different parts such as background technology, technical field, technical solution, and embodiments. If the user selects technical solution for retrieval, the vectors in the technical solution section can be directly compared, ignoring other vectors. If the user does not specify a particular vector, the comparison results of different vectors can be calculated separately and then weighted averaged. The weights for the technical solution and embodiment sections can be set higher, such as a weight of 1 for background technology, 1.5 for technical field, 2 for technical solution, and 1.5 for embodiment. Users can also adjust these weights.

[0112] Finally, in step 106 above, as Figure 2 The "fusion search" shown:

[0113] There are two ways: (1) Use vector comparison to retrieve the N most similar results in the database, and then use the Boolean constraints obtained in step 1 to filter and select the N most similar results. (2) First use the Boolean constraints obtained in step 1 to filter and select the N most similar results in the result set.

[0114] As can be seen from the above, in one embodiment, the semantic retrieval result obtained by fusing the retrieval based on the relationship between the vector model and the updated atomic conditions may include:

[0115] Based on the vector model, the most similar results are retrieved from a pre-established database using a vector comparison method;

[0116] By updating the relationships between atomic conditions, the most similar results are filtered and selected to obtain the final semantic retrieval results.

[0117] As can be seen from the above, in one embodiment, the semantic retrieval result obtained by fusing the retrieval based on the relationship between the vector model and the updated atomic conditions may include:

[0118] Based on the vector model, the relationships between the updated atomic conditions are filtered and selected to obtain preliminary search results;

[0119] In the initial search results, the most similar results are sorted by vector comparison method and taken as the final semantic search results.

[0120] Combine the Boolean expression obtained in the above steps to form a complete search expression: CA=Country A AND (TI=soybean milk maker OR pressure reliever) AND PT=invention AND RS=TOP(10) AND IPC=(A47J31% OR A23C11%) AND R=(VectorToSearch).

[0121] Searching is performed using the expression generated in the above steps, and the top 10 results are taken. The list and search results are shown in Table 4 below:

[0122] Table 4: Search Results

[0123]

[0124] Through comparison, the TOP10 patent texts in the search results are highly similar to the text of the patent application to be detected, which verifies the accuracy of the search results.

[0125] In summary, the beneficial technical effects that can be achieved by the semantic retrieval method provided in the embodiments of the present invention are:

[0126] 1. More accurate semantic understanding: Through double-layer parsing and chain-of-thought reasoning, errors in condition recognition and logical combination are reduced.

[0127] 2. Fine-grained semantic capture is achieved, and the anti-noise ability is stronger: Through entity recognition and establishment of entity relationships, the deficiency of ignoring details of technical solutions caused by only using keywords is fundamentally avoided.

[0128] 3. Strong scalability: It supports adjustable weight and tuple graph expansion, and can adapt to different fields and different user requirements.

[0129] An embodiment of the present invention also provides a semantic retrieval device, as described in the following embodiments. Since the problem-solving principle of the device is similar to that of the semantic retrieval method, the implementation of the device can refer to the implementation of the semantic retrieval method, and repetitions will not be described again.

[0130] Figure 4 is a schematic structural diagram of the semantic retrieval device in an embodiment of the present invention, as Figure 4 shows, the device comprises:

[0131] an atomic condition information table generation unit 01, configured to input search content input by a user into an atomic condition recognition model, extract key entity information, and generate an atomic condition information table, where each key entity information corresponds to one atomic condition; the atomic condition recognition model is pre-trained and generated according to sample data between historical search content and historical key entity information;

[0132] Atomic classification unit 02 is used to input the semantic information of the user-input search content into an external general model to obtain the expression pseudocode corresponding to the user-input search content; it also compiles the atomic conditions in the atomic condition information table, the user-input search content, and the expression pseudocode into prompt content and inputs it into the external general model for classification to obtain atomic conditions with classification types; the external general model is pre-trained and generated based on the sample data of the relationship between the search content and the expression pseudocode, and the sample data of the relationship between the prompt content and the atomic conditions with classification types.

[0133] Atomic assembly unit 03 is used to extract and assemble atomic conditions with classification types according to the semantic information of the user input search content to obtain the relationship between atomic conditions;

[0134] The element-limited fusion unit 04 is used to extract search elements from the semantic information of the search content input by the user, expand the search elements into entities, obtain the search entity element classification limitation conditions according to the classification of the expanded search entity elements, and merge them into the atomic condition information table to obtain atomic conditions with classification type again, so as to re-extract and combine to obtain the relationship between updated atomic conditions.

[0135] Vector model building unit 05 is used to build a vector model from the expanded retrieved entity elements according to different application needs, using triplet, quadruple, or multi-tuple models. Triplets are used to record the relationship between the subject and the object, quadruplets are used to record the relationship between the subject, subject attributes, and object, and multi-tuples are used to record the relationship between the subject, subject attributes, object, and object attributes, as well as the conditions required to generate the relationship.

[0136] The fusion retrieval unit 06 is used to obtain semantic retrieval results by fusion retrieval based on the relationship between the vector model and the updated atomic conditions.

[0137] In one embodiment, the atomic conditions in the atomic condition information table, the user input search content, and the expression pseudocode rules are compiled together into prompt content and sent to an external general large model for classification to obtain atomic conditions with classification types, including the following three types of atomic conditions: target restriction conditions, content feature conditions, and metadata filtering conditions.

[0138] In one embodiment, the element-defining fusion unit is specifically used for:

[0139] If the atomic condition information contains similar patent limitation conditions, then the classification number of the similar patent limitation conditions is obtained and used as the classification limitation condition for the search entity element, and merged into the atomic condition information table;

[0140] If the atomic condition information contains only keywords, then by calculating the keyword vector, comparing it with the pre-vectorized classification table through an external classification query tool, and returning the N classifications with the highest vector similarity as the classification criteria for the retrieved entity elements, they are merged into the atomic condition information table.

[0141] If the atomic condition information does not contain either similar patent limitation conditions or keywords, then the classification limitation conditions are ignored.

[0142] In one embodiment, the semantic retrieval method described above may further include a weighting unit, used to: assign different weights to triples, quadruples or multi-tuples, or assign different weights to different elements in triples, quadruples or multi-tuples.

[0143] In one embodiment, the fusion retrieval unit is specifically used for:

[0144] Based on the vector model, the most similar results are retrieved from a pre-established database using a vector comparison method;

[0145] By updating the relationships between atomic conditions, the most similar results are filtered and selected to obtain the final semantic retrieval results.

[0146] In one embodiment, the fusion retrieval unit is specifically used for:

[0147] Based on the vector model, the relationships between the updated atomic conditions are filtered and selected to obtain preliminary search results;

[0148] In the initial search results, the most similar results are sorted by vector comparison method and taken as the final semantic search results.

[0149] Based on the aforementioned inventive concept, such as Figure 5 As shown, the present invention also proposes a computer device 500, including a memory 510, a processor 520, and a computer program 530 stored in the memory 510 and executable on the processor 520. When the processor 520 executes the computer program 530, it implements the aforementioned semantic retrieval method.

[0150] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the semantic retrieval method described above.

[0151] This invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the semantic retrieval method described above.

[0152] The beneficial technical effects of the semantic retrieval scheme provided in this embodiment of the invention are:

[0153] First, in this embodiment of the invention, an atomic condition recognition model is pre-trained based on sample data between historical search content and historical key entity information. When performing real-time semantic retrieval, the user input search content is input into the pre-trained atomic condition recognition model to extract key entity information and generate an atomic condition information table, with each key entity information corresponding to an atomic condition.

[0154] Secondly, this embodiment of the invention also utilizes an external general-purpose large model pre-trained based on sample data of the relationship between the search content and the expression pseudocode, and sample data of the relationship between the prompt content and the atomic conditions with classification types. During real-time semantic retrieval, the semantic information of the user-input search content is input into the external general-purpose large model to obtain the expression pseudocode corresponding to the user-input search content; the atomic conditions in the atomic condition information table, the user-input search content, and the expression pseudocode are arranged together as prompt content and input into the external general-purpose large model for classification to obtain atomic conditions with classification types.

[0155] Based on the above, this embodiment of the invention proposes a semantic retrieval scheme that combines small models and large models.

[0156] Furthermore, in this embodiment of the invention, atomic conditions with classification types are extracted and combined according to the semantic information of the user-input search content to obtain the relationship between atomic conditions; search elements are extracted from the semantic information of the user-input search content, and entity expansion is performed on the search elements; based on the classification of the expanded search entity elements, search entity element classification constraints are obtained, which are then merged into the atomic condition information table to obtain atomic conditions with classification types again, so as to re-extract and combine to obtain the updated relationship between atomic conditions. This realizes the fusion of search entity element classification constraints and atomic condition information table, improves the accuracy of the relationship between atomic conditions, and thus improves the accuracy of subsequent semantic retrieval.

[0157] Furthermore, according to different application requirements, the extended retrieval entity elements in this embodiment of the invention establish a vector model composed of triples, quadruples, or multi-tuples. Triples are used to record the relationship between the subject and the object, quadruples are used to record the relationship between the subject, subject attributes, and the object, and multi-tuples are used to record the relationship between the subject, subject attributes, the object, and object attributes, as well as the conditions required to generate the relationship. The vector model composed of triples, quadruples, or multi-tuples can improve the accuracy of subsequent semantic retrieval.

[0158] Finally, based on the relationship between the vector model and the updated atomic conditions, the present invention re-integrates the retrieval to obtain semantic retrieval results, thus realizing a semantic retrieval optimization method based on thought chain reasoning.

[0159] In summary, this invention proposes a semantic retrieval optimization method that combines small and large models and is based on thought chain reasoning, which can achieve efficient and accurate semantic retrieval in large-scale data environments.

[0160] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0161] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0162] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0163] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0164] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A semantic retrieval method, characterized in that, include: The user inputs the search content into the atomic condition recognition model, extracts key entity information, and generates an atomic condition information table. Each key entity information corresponds to an atomic condition, and the atomic condition is the system field and content fragment corresponding to each key entity information. The atomic condition recognition model is pre-trained based on sample data between historical search content and historical key entity information. The semantic information of the user's search query is input into an external general model to obtain the pseudocode of the expression corresponding to the user's search query. The atomic conditions in the atomic condition information table, the user-input search content, and the expression pseudocode are arranged together as prompts and input into an external general model for classification to obtain atomic conditions with classification types. The external general model is pre-trained and generated based on the sample data of the relationship between the search content and the expression pseudocode, as well as the sample data of the relationship between the prompts and the atomic conditions with classification types. The relationships between atomic conditions are obtained by extracting and combining the semantic information of the user-input search content based on the atomic conditions with classification types. Extract search elements from the semantic information of the search content entered by the user, and expand the search elements into entities; Based on the expanded retrieval entity element classification, the retrieval entity element classification constraints are obtained, which are then integrated into the atomic condition information table to obtain atomic conditions with classification types again, so as to re-extract and combine them to obtain the updated relationships between atomic conditions. Based on different application needs, the expanded retrieval entity elements are used to establish a vector model consisting of triples, quadruples, or multi-tuples. Triples are used to record the relationship between the subject and the object, quadruples are used to record the relationship between the subject, subject attributes, and the object, and multi-tuples are used to record the relationship between the subject, subject attributes, the object, and the object attributes, as well as the conditions required to generate the relationship. Based on the relationship between the vector model and the updated atomic conditions, the semantic retrieval results are obtained through fusion retrieval.

2. The method as described in claim 1, characterized in that, The atomic conditions in the atomic condition information table, the user input search content, and the expression pseudocode rules are compiled together into prompt content and sent to an external general model for classification to obtain atomic conditions with classification types, including the following three types of atomic conditions: target restriction conditions, content feature conditions, and metadata filtering conditions.

3. The method as described in claim 1, characterized in that, Based on the expanded retrieval entity element classification, the retrieval entity element classification constraints are obtained, and these are merged into the atomic condition information table to re-obtain atomic conditions with classification types. This allows for the re-extraction and combination of updated atomic conditions, including: If the atomic condition information contains similar patent limitation conditions, then the classification number of the similar patent limitation conditions is obtained and used as the classification limitation condition for the search entity element, and merged into the atomic condition information table; If the atomic condition information contains only keywords, then by calculating the keyword vector, comparing it with the pre-vectorized classification table through an external classification query tool, and returning the N classifications with the highest vector similarity as the classification criteria for the retrieved entity elements, they are merged into the atomic condition information table. If the atomic condition information does not contain either similar patent limitation conditions or keywords, then the classification limitation conditions are ignored.

4. The method as described in claim 1, characterized in that, Also includes: Assign different weights to triples, quadruples, or multi-tuples, or assign different weights to different elements in triples, quadruples, or multi-tuples.

5. The method as described in claim 1, characterized in that, Based on the relationship between the vector model and the updated atomic conditions, the fusion retrieval yields semantic retrieval results, including: Based on the vector model, the most similar results are retrieved from a pre-established database using a vector comparison method; By updating the relationships between atomic conditions, the most similar results are filtered and selected to obtain the final semantic retrieval results.

6. The method as described in claim 1, characterized in that, Based on the relationship between the vector model and the updated atomic conditions, the fusion retrieval yields semantic retrieval results, including: Based on the vector model, the relationships between the updated atomic conditions are filtered and selected to obtain preliminary search results; In the initial search results, the most similar results are sorted by vector comparison method and taken as the final semantic search results.

7. A semantic retrieval device, characterized in that, include: The atomic condition information table generation unit is used to input the user's search content into the atomic condition recognition model, extract key entity information, and generate an atomic condition information table. Each key entity information corresponds to an atomic condition, and the atomic condition is the system field and content fragment corresponding to each key entity information. The atomic condition recognition model is pre-trained and generated based on sample data between historical search content and historical key entity information. Atomic fractal units are used to input the semantic information of the user's input search content into an external general large model to obtain the expression pseudocode corresponding to the user's input search content; The atomic conditions in the atomic condition information table, the user-input search content, and the expression pseudocode are arranged together as prompts and input into an external general model for classification to obtain atomic conditions with classification types. The external general model is pre-trained and generated based on the sample data of the relationship between the search content and the expression pseudocode, as well as the sample data of the relationship between the prompts and the atomic conditions with classification types. Atomic assembly unit is used to extract and assemble atomic conditions with classification types according to the semantic information of the user input search content to obtain the relationship between atomic conditions; The element-limited fusion unit is used to extract search elements from the semantic information of the search content input by the user and to expand the search elements into entities. Based on the expanded retrieval entity element classification, the retrieval entity element classification constraints are obtained, which are then integrated into the atomic condition information table to obtain atomic conditions with classification types again, so as to re-extract and combine them to obtain the updated relationships between atomic conditions. The vector model building unit is used to build a vector model from the expanded retrieved entity elements according to different application requirements, using triplet, quadruple, or multi-tuple models. Triplets are used to record the relationship between the subject and the object, quadruplets are used to record the relationship between the subject, subject attributes, and object, and multi-tuples are used to record the relationship between the subject, subject attributes, object, and object attributes, as well as the conditions required to generate the relationship. The fusion retrieval unit is used to obtain semantic retrieval results by fusing retrieval based on the relationship between the vector model and the updated atomic conditions.

8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1 to 6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Pressure relief device of soybean milk machine

    CN102397011A

  • Intelligent patent similarity searching method and device based on semantic retrieval

    CN113821646A

  • Scientific and technological novelty retrieval method and device based on semantic elements, equipment and medium

    CN119669389A