Knowledge retrieval and matching method and system based on large model driving
Through the knowledge retrieval and matching method driven by a large model, by updating vocabulary classification and stimulating the analysis of key extracted words and relationship extracted words, the problem of low matching between knowledge retrieval results and user intentions is solved, and efficient and accurate knowledge retrieval and matching is achieved.
Patent Information
- Application Number
- CN202510667880.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-05
AI Technical Summary
Existing technologies fail to effectively combine semantic context and historical usage data, resulting in a low match between knowledge retrieval results and user intent. They also lack deep modeling of dynamic semantic relationships between knowledge vocabulary, limiting semantic recommendation and associative reasoning capabilities.
Through a large model-driven approach, we update vocabulary classification, calculate inverse document frequency and call counts, stimulate key extraction words, set trigger values to filter words, analyze the stimulation time ratio and utilization rate of relational extraction words, and perform relational chain feature analysis to achieve accurate and efficient knowledge retrieval and matching.
It improves the accuracy and pertinence of knowledge retrieval, reduces the cost of prompt words, and improves the precision of prompts.
Smart Images

Figure CN120596513A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of retrieval and matching technology, and more specifically, to a knowledge retrieval and matching method and system driven by a large model. Background Art
[0002] With the development of information technology and the widespread application of artificial intelligence, knowledge management and information retrieval have become core needs in fields such as education, healthcare, law, and scientific research. Especially in the big data environment, users face the redundant interference of massive amounts of unstructured information. With the continuous accumulation of user interactions, modeling vocabulary usage frequency, retrieval paths, and historical call history based on dynamic learning mechanisms also provides important support for the structured classification and semantic matching of knowledge vocabulary.
[0003] At present, due to the failure to combine semantic context and historical usage data, the match between retrieval results and user intentions is low, resulting in the inability to effectively identify and classify the semantic drift of the same word in different contexts, vague attribution of knowledge vocabulary, and a lack of deep modeling of the dynamic semantic relationship between knowledge vocabulary, which limits the ability of semantic recommendation and associative reasoning. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a knowledge retrieval and matching method and system driven by a large model, which dynamically updates vocabulary classification and inverse document frequency calculation, combines vocabulary coverage and retrieval frequency to stimulate the key extraction word mechanism, sets trigger values to filter and optimize key extraction words, and uses the activation time ratio and the historical retrieval frequency and collection increment of the relationship extraction words to calculate the utilization rate, and then performs relationship chain feature analysis and provides corrections, thereby achieving accurate and efficient knowledge retrieval and matching to solve the problems raised in the above-mentioned background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] The knowledge retrieval and matching method based on large model driving includes the following steps:
[0007] Step S1: before the user searches for knowledge words, update the vocabulary classification, count the number of documents under each preset topic and the number of documents containing the knowledge words in the corresponding topic, calculate the inverse document frequency of the knowledge words, and collect the number of times the knowledge words are called in each preset topic;
[0008] Step S2: Calculate the vocabulary attribution characteristics of the knowledge vocabulary by integrating the inverse document frequency of the knowledge vocabulary and the number of calls in the preset topic, and select the knowledge vocabulary according to the vocabulary attribution characteristics to be included in the preset topic;
[0009] Step S3: When the user searches for knowledge vocabulary, the vocabulary coverage and search frequency of the knowledge vocabulary are retrieved to determine whether to trigger key extraction words. When activating key extraction words, a trigger value is set to filter words under the corresponding topic of the knowledge vocabulary for provision. The activation time is set to calculate the activation time ratio of each relation extraction word.
[0010] Step S4: Call the historical search frequency and collection increment of each relational extraction word in the historical time and calculate the utilization rate of the relational extraction word, analyze the relationship chain characteristics of each relational extraction word and the current knowledge vocabulary based on the comprehensive triggering time ratio and utilization rate, and provide corrections for the relational extraction words of the current knowledge vocabulary based on the comprehensive trigger value and relationship chain characteristics.
[0011] In a preferred embodiment, in step S1, the preset subject is the default classification preset before the knowledge vocabulary is updated, and the subject is the category of the knowledge vocabulary. The vocabulary is updated before the user searches for the knowledge vocabulary.
[0012] In a preferred embodiment, in step S1, a knowledge vocabulary search is performed under each preset topic, the number of documents in each preset topic and the number of documents in each knowledge vocabulary in the corresponding topic are counted, and the inverse document frequency of each knowledge vocabulary is calculated: IDF(a, D)=log(|D| / |D a |), where IDF(a, D) is the inverse document frequency of knowledge words in each preset topic, a is the knowledge word, D is the total number of documents in each preset topic, and D a is the number of documents containing knowledge term a in each preset topic.
[0013] In a preferred embodiment, in step S2, the inverse document frequency and call count of the comprehensive knowledge vocabulary in each preset topic are used to calculate the vocabulary attribution feature using a logistic regression algorithm;
[0014] For the same knowledge vocabulary, the values of its vocabulary attribution characteristics in each preset topic are compared. Among the preset topics, the preset topic with the largest value of the vocabulary attribution characteristics of the knowledge vocabulary is selected as the classification topic after the corresponding knowledge vocabulary is updated.
[0015] In a preferred embodiment, in step S3, by setting a unique attribution mapping between knowledge vocabulary and preset topics, the number of documents containing the corresponding vocabulary is counted within the corresponding topic, and the ratio is calculated with the total number of documents in the preset topic to obtain the vocabulary coverage of the knowledge vocabulary;
[0016] Record the search log, accumulate the total number of times the knowledge vocabulary is searched within the set time window, and compare it with the set time window length to obtain the search frequency of the knowledge vocabulary;
[0017] The vocabulary coverage and retrieval frequency of knowledge vocabulary are standardized and substituted into the geometric mean method to calculate the activation coefficient.
[0018] In a preferred embodiment, in step S3, the excitation coefficient is compared with a preset excitation threshold. If the excitation coefficient is greater than or equal to the excitation threshold, the key extraction word is excited. If the excitation coefficient is less than the excitation threshold, the key extraction word operation is not performed.
[0019] The key words to be stimulated are the words that are associated and related when users search for knowledge words;
[0020] To set the trigger value when stimulating key extraction words, enter the trigger value provision mechanism. The specific steps are as follows:
[0021] Assign a value to each key extraction word, and filter it according to the trigger value size corresponding to each key extraction word and the preset number of provided words, and provide the filtered key extraction words;
[0022] The trigger values corresponding to the key extraction words are sorted in order from large to small, and the key extraction words consistent with the preset number are provided according to the preset number.
[0023] In a preferred embodiment, in step S3, after providing the key extraction words, the activation time is set and the activation time ratio of each relation extraction word is calculated;
[0024] After stimulating the key extraction words, set the stimulation time, detect the length of stay on each relation extraction word page respectively, and calculate the ratio with the total length of stay on all relation extraction word pages to obtain the stimulation time ratio of each relation extraction word;
[0025] The historical search frequency and collection increment of each relation extraction word in the historical time are called to obtain the utilization rate of the relation extraction word.
[0026] In a preferred embodiment, in step S4, within a set historical time window, the historical search frequency and collection increment of the relational extraction words are counted respectively, and after normalization, they are integrated using the weighted average method to obtain the utilization rate of the relational extraction words;
[0027] The activation time ratio of each relation extraction word and the utilization rate of the relation extraction word are standardized and substituted into the polynomial regression calculation to obtain the relation chain feature;
[0028] If the relation extraction word has a relation chain feature, the corresponding trigger value will be invalidated, and the relation chain feature will be substituted into the corresponding relation extraction word for correction.
[0029] In a preferred embodiment, in step S4, the following rules are set when searching the knowledge vocabulary next time:
[0030] Rule 1: When a relational extraction word has a relational chain feature and the number of relational extraction words with the relational chain feature exceeds the preset number of provided words, the relationship chain feature size is compared to select the preset number of relational extraction words for provision;
[0031] Rule 2: When the relation extraction words have chain features and the number of relation extraction words with chain features is lower than the preset number, after the relation extraction words with chain features are provided, the difference below the preset number is determined and the trigger value provision mechanism is entered.
[0032] A knowledge retrieval and matching system driven by a large model is used to implement the above-mentioned knowledge retrieval and matching method driven by a large model, including a vocabulary update module, an attribution analysis module, an excitation extraction module, and a relationship correction module, with electrical signal connections between the modules;
[0033] The vocabulary update module is used to update vocabulary classification before users search for knowledge vocabulary, collect the inverse document frequency of each knowledge vocabulary and the number of times each knowledge vocabulary is called in each preset topic;
[0034] The attribution analysis module is used to calculate the vocabulary attribution characteristics of knowledge vocabulary by comprehensively analyzing the inverse document frequency of knowledge vocabulary and the number of calls in the preset topic. Based on the vocabulary attribution characteristics, knowledge vocabulary is selected and included in the preset topic. Knowledge vocabulary that is not included in the preset topic is screened out and updated in vocabulary classification through the tag inclusion method.
[0035] The trigger extraction module is used to retrieve the vocabulary coverage and search frequency of knowledge vocabulary when the user searches for knowledge vocabulary to determine whether to trigger key extraction words. When stimulating key extraction words, the trigger value is set to filter the words under the corresponding topic of knowledge vocabulary for provision, and the trigger time is set to calculate the trigger time ratio of each relation extraction word.
[0036] The relationship correction module is used to call the historical search frequency and collection increment of each relationship extraction word in the historical time and calculate the utilization rate of the relationship extraction word. It comprehensively analyzes the relationship chain characteristics of each relationship extraction word and the current knowledge vocabulary based on the comprehensive triggering time ratio and utilization rate, and provides corrections for the relationship extraction words of the current knowledge vocabulary based on the comprehensive trigger value and relationship chain characteristics.
[0037] The technical effects and advantages of the knowledge retrieval and matching method and system driven by a large model of the present invention are as follows:
[0038] The present invention updates vocabulary classification before the user searches for knowledge vocabulary, counts the number of documents under each preset topic and the number of documents containing the knowledge vocabulary in the corresponding topic, calculates the inverse document frequency of the knowledge vocabulary, collects the number of times the knowledge vocabulary is called in each preset topic, calculates the vocabulary attribution characteristics of the knowledge vocabulary, selects the knowledge vocabulary and includes it in the preset topic according to the vocabulary attribution characteristics, and when the user searches for knowledge vocabulary, retrieves the vocabulary coverage and detection frequency of the knowledge vocabulary to judge whether to trigger key extraction words, improves the accuracy of the search and sets the trigger value pertinently, filters the vocabulary under the corresponding topic of the knowledge vocabulary, calculates the activation time ratio, calls the historical search frequency and collection increment of each relation extraction word in the historical time, calculates the utilization rate of the relation extraction word, analyzes the relationship chain characteristics between each relation extraction word and the current knowledge vocabulary, and provides corrections for the relation extraction words of the current knowledge vocabulary, thereby reducing the prompt word cost while improving the prompt accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 This is a flow chart of the knowledge retrieval and matching method driven by a large model according to the present invention.
[0040] Figure 2 Schematic diagram of the knowledge retrieval and matching system driven by a large model according to the present invention. DETAILED DESCRIPTION
[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0042] The present invention updates vocabulary classification before the user retrieves knowledge vocabulary, counts the number of documents under each preset topic and the number of documents where the knowledge vocabulary is located in the corresponding topic, calculates the inverse document frequency of the knowledge vocabulary, collects the number of times the knowledge vocabulary is called in each preset topic, calculates the vocabulary attribution characteristics of the knowledge vocabulary, selects the knowledge vocabulary and includes it in the preset topic according to the vocabulary attribution characteristics, and when the user retrieves the knowledge vocabulary, retrieves the vocabulary coverage and detection frequency of the knowledge vocabulary to determine whether to trigger key extraction words, sets the trigger value to filter the vocabulary under the corresponding topic of the knowledge vocabulary for provision, calculates the activation time ratio, calls the historical search frequency and collection increment of each relation extraction word in the historical time, calculates the utilization rate of the relation extraction word, analyzes the relationship chain characteristics between each relation extraction word and the current knowledge vocabulary, and provides corrections for the relation extraction words of the current knowledge vocabulary, thereby reducing the prompt word cost while improving the prompt accuracy.
[0043] Example 1, a knowledge retrieval and matching method driven by a large model, such as Figure 1 As shown, the following steps are included:
[0044] Step S1: before the user searches for knowledge words, update the vocabulary classification, count the number of documents under each preset topic and the number of documents containing the knowledge words in the corresponding topic, calculate the inverse document frequency of the knowledge words, and collect the number of times the knowledge words are called in each preset topic;
[0045] Step S2: Calculate the vocabulary attribution characteristics of the knowledge vocabulary by integrating the inverse document frequency of the knowledge vocabulary and the number of calls in the preset topic, and select the knowledge vocabulary according to the vocabulary attribution characteristics to be included in the preset topic;
[0046] Step S3: When the user searches for knowledge vocabulary, the vocabulary coverage and search frequency of the knowledge vocabulary are retrieved to determine whether to trigger key extraction words. When activating key extraction words, a trigger value is set to filter words under the corresponding topic of the knowledge vocabulary for provision. The activation time is set to calculate the activation time ratio of each relation extraction word.
[0047] Step S4: Call the historical search frequency and collection increment of each relational extraction word in the historical time and calculate the utilization rate of the relational extraction word, analyze the relationship chain characteristics of each relational extraction word and the current knowledge vocabulary based on the comprehensive triggering time ratio and utilization rate, and provide corrections for the relational extraction words of the current knowledge vocabulary based on the comprehensive trigger value and relationship chain characteristics.
[0048] The specific implementation is as follows:
[0049] In step S1, the preset subject is the default classification preset before the knowledge vocabulary is updated. The subject is the category of the knowledge vocabulary, and the vocabulary is updated before the user searches for the knowledge vocabulary.
[0050] Call the knowledge vocabulary in the document database, perform knowledge vocabulary search under each preset topic, count the number of documents in each preset topic and the number of documents where each knowledge vocabulary is located in the corresponding topic, and calculate the inverse document frequency of each knowledge vocabulary: IDF(a, D) = log(|D| / |D a |), where IDF(a, D) is the inverse document frequency of knowledge words in each preset topic, a is the knowledge word, D is the total number of documents in each preset topic, and D a is the number of documents containing knowledge vocabulary a in each preset topic;
[0051] Inverse document frequency is used to measure the importance of knowledge vocabulary in the corresponding subject document set. The lower the inverse document frequency, the more common the corresponding knowledge vocabulary is in the subject document set, and the more it should be included in the corresponding subject.
[0052] For the knowledge vocabulary in the document database, a period of time is selected as the collection time, and the number of times the knowledge vocabulary is called in each preset topic is collected within the collection time.
[0053] It should be noted that a document database is a data set that stores and manages document formats, and is used to call knowledge vocabulary in this example.
[0054] In step S2, the inverse document frequency and call count of the comprehensive knowledge vocabulary in each preset topic are used to calculate the vocabulary attribution feature using the logistic regression algorithm: L = 1 / 1 + e -z , where e is the natural base, z is the intermediate parameter and z is the difference between the number of calls of the knowledge vocabulary in each preset topic and the inverse document frequency, and L is the vocabulary attribution feature of the knowledge vocabulary in each preset topic;
[0055] For the same knowledge vocabulary, the values of its vocabulary attribution characteristics in each preset topic are compared. Among the preset topics, the preset topic with the largest value of the vocabulary attribution characteristics of the knowledge vocabulary is selected as the classification topic after the corresponding knowledge vocabulary is updated.
[0056] In step S3, after the user enters the search term, the key word extraction mechanism is determined based on the document data in the preset topic corresponding to the current term;
[0057] After the user enters a certain knowledge word, the distribution of documents in the preset topic that the word corresponds to will determine whether it will trigger the key word extraction mechanism. This judgment is based on the vocabulary coverage and search frequency of the knowledge word.
[0058] The vocabulary coverage of a knowledge vocabulary is the degree of document coverage of the corresponding knowledge vocabulary retrieved by the user in the preset topic to which it belongs, that is, the breadth of its distribution in all documents under the topic. The acquisition logic is to set a unique attribution mapping between the knowledge vocabulary and the preset topic, count the number of documents containing the corresponding vocabulary in the corresponding topic, and calculate the ratio with the total number of documents in the preset topic to obtain the vocabulary coverage of the knowledge vocabulary;
[0059] It should be noted that the preset topics in this method are fixed topic sets constructed by the system based on domain knowledge or model clustering training results, and each knowledge word belongs to only one topic. In the case of low coverage values, the system can improve the completeness of the search results by introducing a related word expansion mechanism, which will not be elaborated here.
[0060] The retrieval frequency of a knowledge word measures the degree of attention paid to the word in the user's historical behavior and reflects its popularity in the retrieval log. Its acquisition logic is to record the retrieval log, accumulate the total number of times the knowledge word is retrieved within a set time window, and compare it with the set time window length to obtain the retrieval frequency of the knowledge word;
[0061] It should be noted that recording search logs is the process of the system automatically capturing, structuring, and timestamping the query terms submitted by users through the search interface. This process is based on unified formatting rules to ensure the consistency of terms and statistical dimensions. The time window is determined by our experimenters based on the system resource carrying capacity and the statistical stability of user behavior, and will not be elaborated here.
[0062] The vocabulary coverage and retrieval frequency of knowledge vocabulary are standardized and substituted into the geometric mean method to calculate the activation coefficient;
[0063] It should be noted that the standardization methods include but are not limited to standard linear transformation based on interval scaling, Z-Score standardization method based on statistics, or normalization method based on nonlinear mapping function. The application methods of standardization are not described in detail here.
[0064] Specifically, the calculation formula of the geometric mean method is:
[0065] In the formula, S is the activation coefficient, A is the vocabulary coverage of knowledge vocabulary, and b is the retrieval frequency of knowledge vocabulary;
[0066] The geometric mean method is a common calculation method for those skilled in the art and is common knowledge, so it will not be described in detail here.
[0067] Compare the excitation coefficient with the preset excitation threshold. If the excitation coefficient is greater than or equal to the excitation threshold, the key word extraction operation is activated. If the excitation coefficient is less than the excitation threshold, the key word extraction operation is not performed.
[0068] It should be noted that the activation threshold was obtained by our experimenters based on the frequency distribution of historical search behavior data and the average penetration of vocabulary coverage in the topic level, which will not be elaborated here;
[0069] Among them, the stimulating key extraction words are the word associations and associated words when users search for knowledge words;
[0070] For example, when a user searches for the term "Three Kingdoms", key words are extracted, and the associated words include but are not limited to "Romance of the Three Kingdoms", "Five Tiger Generals", "Five Sons of the Army", etc.
[0071] To set the trigger value when stimulating key extraction words, enter the trigger value provision mechanism. The specific steps are as follows:
[0072] Assign a value to each key extraction word, and filter it according to the trigger value size corresponding to each key extraction word and the preset number of provided words, and provide the filtered key extraction words;
[0073] Furthermore, the trigger values corresponding to the key extraction words are sorted in order from large to small, and according to the preset number of provided words, the key extraction words that are consistent with the preset number of provided words are provided.
[0074] It should be noted that the trigger values corresponding to each key extracted word are randomly generated by our experimenters. The random generation mechanism is based on the natural language context perturbation characteristics and the sparse distribution structure modeling results of the semantic space, which will not be elaborated here.
[0075] Furthermore, the preset number of provided keywords is the number of key words that the system outputs at one time, which is used to control the scale of the recommendation results. The specific value is obtained by our experimenters based on typical user behavior log analysis and keyword acceptance survey data, so it will not be detailed here.
[0076] For example, when a user searches for the term "Three Kingdoms", the key extraction words are triggered, and the association and related words include "Romance of the Three Kingdoms", "Five Tiger Generals", "Five Sons of Good Generals" and "Lü Bu". The trigger value of "Romance of the Three Kingdoms" is randomly preset to be 0.45, the trigger value of "Five Tiger Generals", the trigger value of "Five Sons of Good Generals" is 0.21, the trigger value of "Lü Bu" is 0.24, and the preset number of provided words is three. Therefore, the key extraction words provided are "Romance of the Three Kingdoms", "Lü Bu" and "Five Tiger Generals".
[0077] After providing the key extraction words, set the activation time and calculate the activation time ratio of each relation extraction word;
[0078] Among them, the activation time ratio of each relational extraction word refers to the attention intensity shown by users on all recommended relational extraction word pages within the set activation time window, that is, the proportion of the user's stay time on the page of a specific relational extraction word in the total activation time. It is used to measure the interactive importance of the relational extraction word in the activation stage. Its acquisition logic is to set the activation time after stimulating the key extraction word, detect the stay time of each relational extraction word page respectively, and calculate the ratio with the total stay time of all relational extraction word pages to obtain the activation time ratio of each relational extraction word;
[0079] The triggering time was obtained by our researchers based on the statistical results of the mean fluctuation of user behavior and the recommendation content interaction continuity judgment model, and will not be elaborated here.
[0080] Furthermore, the dwell time on each relational word extraction page refers to the continuous duration recorded from the time a user opens the corresponding extracted word page to the time the user actively closes the page, jumps to another page, or times out due to inactivity, in seconds or milliseconds. This dwell time is determined by the system through a comprehensive assessment of browser behavior, foreground activity, and clickstream data recorded by the user-side interaction detection module, and is not limited here.
[0081] In step S4, the historical search frequency and collection increment of each relation extraction word in the historical time are called to obtain the utilization rate of the relation extraction word;
[0082] The utilization rate of a relational word refers to the comprehensive use value of the relational word in the actual user interaction behavior within a set historical time window, reflecting its long-term access and retention activity. It is the result indicator after combining the historical search frequency and the collection increment. Its acquisition logic is to calculate the historical search frequency and collection increment of the relational word within the set historical time window, perform normalization, and then use the weighted average method to integrate them to obtain the utilization rate of the relational word.
[0083] The historical search frequency of the relational extracted word is calculated by counting the cumulative number of times users searched for the extracted word within a set historical time window, and the collection increment is calculated by counting the number of times the extracted word was newly added to the collection within the historical time window. We will not elaborate on this here.
[0084] The historical time window was set by our experimenters based on the clustering results of user usage cycle behavior and the statistical distribution curve of search cycle, so we will not elaborate on it here.
[0085] Furthermore, collection behavior refers to the data records generated by users when reading the extracted word content page by clicking on the collection, adding to favorites, or setting as a frequently used tag. The system extracts and archives them through the user operation log and account behavior tracking module;
[0086] The activation time ratio of each relation extraction word and the utilization rate of the relation extraction word are standardized and substituted into the polynomial regression calculation to obtain the relation chain feature;
[0087] The standardization process has been described in the above content and will not be repeated here;
[0088] The polynomial regression calculation formula is as follows:
[0089]
[0090] Where R i is the relationship chain feature, S i is the activation time ratio of each relation extraction word, F iis the utilization rate of relational extracted words, c is a constant term, γ1 and γ2 are the weights of the activation time ratio of each relational extracted word and the utilization rate of the relational extracted word, respectively. The specific weight values are obtained by the experimenters based on the sample fitting goodness of fit and the principle of minimizing the sum of squared errors, and will not be elaborated here.
[0091] If the relation extraction word has a relation chain feature, the corresponding trigger value will be invalidated, and the relation chain feature will be substituted into the corresponding relation extraction word for correction;
[0092] Invalidation processing refers to the process of reducing or setting the weight of the original trigger value to zero after detecting significant relationship chain features in the relationship extraction words to avoid redundancy or information bias in subsequent recommendations caused by the same association logic.
[0093] Furthermore, the next time the knowledge vocabulary is searched, the following rules are set:
[0094] Rule 1: When a relational extraction word has a relational chain feature and the number of relational extraction words with the relational chain feature exceeds the preset number of provided words, the relationship chain feature size is compared to select the preset number of relational extraction words for provision;
[0095] Rule 2: When the relation extraction words have chain features and the number of relation extraction words with chain features is lower than the preset number, after the relation extraction words with chain features are provided, the difference below the preset number is determined and the trigger value provision mechanism is entered.
[0096] Example 2, based on a knowledge retrieval and matching system driven by a large model, such as Figure 2 As shown, it includes a vocabulary updating module, a belonging analysis module, a stimulus extraction module and a relationship correction module, and the modules are connected by electrical signals;
[0097] The functions of each module are as follows:
[0098] The vocabulary update module is used to update vocabulary classification before users search for knowledge vocabulary, collect the inverse document frequency of each knowledge vocabulary and the number of times each knowledge vocabulary is called in each preset topic;
[0099] The attribution analysis module is used to calculate the vocabulary attribution characteristics of knowledge vocabulary by comprehensively analyzing the inverse document frequency of knowledge vocabulary and the number of calls in the preset topic. Based on the vocabulary attribution characteristics, knowledge vocabulary is selected and included in the preset topic. Knowledge vocabulary that is not included in the preset topic is screened out and updated in vocabulary classification through the tag inclusion method.
[0100] The trigger extraction module is used to retrieve the vocabulary coverage and search frequency of knowledge vocabulary when the user searches for knowledge vocabulary to determine whether to trigger key extraction words. When stimulating key extraction words, the trigger value is set to filter the words under the corresponding topic of knowledge vocabulary for provision, and the trigger time is set to calculate the trigger time ratio of each relation extraction word.
[0101] The relationship correction module is used to call the historical search frequency and collection increment of each relationship extraction word in the historical time and calculate the utilization rate of the relationship extraction word. It comprehensively analyzes the relationship chain characteristics of each relationship extraction word and the current knowledge vocabulary based on the comprehensive triggering time ratio and utilization rate, and provides corrections for the relationship extraction words of the current knowledge vocabulary based on the comprehensive trigger value and relationship chain characteristics.
[0102] The above embodiments may be implemented in whole or in part through software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product.
[0103] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application of the technical solution and the invention constraints. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0104] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0105] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0106] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A knowledge retrieval and matching method driven by a large model, characterized by: The following steps are included: Step S1: before the user searches for knowledge words, update the vocabulary classification, count the number of documents under each preset topic and the number of documents containing the knowledge words in the corresponding topic, calculate the inverse document frequency of the knowledge words, and collect the number of times the knowledge words are called in each preset topic; Step S2: Calculate the vocabulary attribution characteristics of the knowledge vocabulary by integrating the inverse document frequency of the knowledge vocabulary and the number of calls in the preset topic, and select the knowledge vocabulary according to the vocabulary attribution characteristics to be included in the preset topic; Step S3: When the user searches for knowledge vocabulary, the vocabulary coverage and search frequency of the knowledge vocabulary are retrieved to determine whether to trigger key extraction words. When activating key extraction words, a trigger value is set to filter words under the corresponding topic of the knowledge vocabulary for provision. The activation time is set to calculate the activation time ratio of each relation extraction word. Step S4: Call the historical search frequency and collection increment of each relational extraction word in the historical time and calculate the utilization rate of the relational extraction word, analyze the relationship chain characteristics of each relational extraction word and the current knowledge vocabulary based on the comprehensive triggering time ratio and utilization rate, and provide corrections for the relational extraction words of the current knowledge vocabulary based on the comprehensive trigger value and relationship chain characteristics.
2. The large model-driven knowledge retrieval and matching method according to claim 1, characterized in that: In step S1, the preset subject is the default classification preset before the knowledge vocabulary is updated. The subject is the category of the knowledge vocabulary, and the vocabulary is updated before the user searches for the knowledge vocabulary.
3. The large model-driven knowledge retrieval and matching method according to claim 1, characterized in that: In step S1, a knowledge vocabulary search is performed under each preset topic, the number of documents in each preset topic and the number of documents in which each knowledge vocabulary is located in the corresponding topic are counted, and the inverse document frequency of each knowledge vocabulary is calculated: IDF(a, D) = log(|D| / |Da|), where IDF(a, D) is the inverse document frequency of the knowledge vocabulary in each preset topic, a is the knowledge vocabulary, D is the total number of documents in each preset topic, and D a is the number of documents containing knowledge term a in each preset topic.
4. The large model-driven knowledge retrieval and matching method according to claim 1 is characterized in that: In step S2, the inverse document frequency and call count of the comprehensive knowledge vocabulary in each preset topic are used to calculate the vocabulary attribution feature using a logistic regression algorithm; For the same knowledge vocabulary, the values of its vocabulary attribution characteristics in each preset topic are compared. Among the preset topics, the preset topic with the largest value of the vocabulary attribution characteristics of the knowledge vocabulary is selected as the classification topic after the corresponding knowledge vocabulary is updated.
5. The large model-driven knowledge retrieval and matching method according to claim 4 is characterized in that: In step S3, by setting a unique attribution mapping between knowledge vocabulary and preset topics, the number of documents containing the corresponding vocabulary is counted within the corresponding topic, and the ratio is calculated with the total number of documents in the preset topic to obtain the vocabulary coverage of the knowledge vocabulary; Record the search log, accumulate the total number of times the knowledge vocabulary is searched within the set time window, and compare it with the set time window length to obtain the search frequency of the knowledge vocabulary; The vocabulary coverage and retrieval frequency of knowledge vocabulary are standardized and substituted into the geometric mean method to calculate the activation coefficient.
6. The large model-driven knowledge retrieval and matching method according to claim 5 is characterized in that: In step S3, the excitation coefficient is compared with a preset excitation threshold. If the excitation coefficient is greater than or equal to the excitation threshold, the key word extraction operation is activated. If the excitation coefficient is less than the excitation threshold, the key word extraction operation is not performed. The key words to be stimulated are the words that are associated and related when users search for knowledge words; To set the trigger value when stimulating key extraction words, enter the trigger value provision mechanism. The specific steps are as follows: Assign a value to each key extraction word, and filter it according to the trigger value size corresponding to each key extraction word and the preset number of provided words, and provide the filtered key extraction words; The trigger values corresponding to the key extraction words are sorted in order from large to small, and the key extraction words consistent with the preset number are provided according to the preset number.
7. The large model-driven knowledge retrieval and matching method according to claim 6, characterized in that: In step S3, after providing the key extraction words, the activation time is set and the activation time ratio of each relation extraction word is calculated; After stimulating the key extraction words, set the stimulation time, detect the length of stay on each relation extraction word page respectively, and calculate the ratio with the total length of stay on all relation extraction word pages to obtain the stimulation time ratio of each relation extraction word; The historical search frequency and collection increment of each relation extraction word in the historical time are called to obtain the utilization rate of the relation extraction word.
8. The large model-driven knowledge retrieval and matching method according to claim 7, characterized in that: In step S4, within the set historical time window, the historical search frequency and collection increment of the relational extraction words are counted respectively, and after normalization, they are integrated using the weighted average method to obtain the utilization rate of the relational extraction words; The activation time ratio of each relation extraction word and the utilization rate of the relation extraction word are standardized and substituted into the polynomial regression calculation to obtain the relation chain feature; If the relation extraction word has a relation chain feature, the corresponding trigger value will be invalidated, and the relation chain feature will be substituted into the corresponding relation extraction word for correction.
9. The large model-driven knowledge retrieval and matching method according to claim 8, characterized in that: In step S4, the following rules are set when searching for the knowledge vocabulary next time: Rule 1: When a relational extraction word has a relational chain feature and the number of relational extraction words with the relational chain feature exceeds the preset number of provided words, the relationship chain feature size is compared to select the preset number of relational extraction words for provision; Rule 2: When the relation extraction words have chain features and the number of relation extraction words with chain features is lower than the preset number, after the relation extraction words with chain features are provided, the difference below the preset number is determined and the trigger value provision mechanism is entered.
10. A knowledge retrieval and matching system driven by a large model, based on the knowledge retrieval and matching method driven by a large model according to any one of claims 1 to 9, characterized in that: It includes a vocabulary updating module, a belonging analysis module, a stimulus extraction module and a relationship correction module, and the modules are connected by electrical signals; The vocabulary update module is used to update vocabulary classification before users search for knowledge vocabulary, collect the inverse document frequency of each knowledge vocabulary and the number of times each knowledge vocabulary is called in each preset topic; The attribution analysis module is used to calculate the vocabulary attribution characteristics of knowledge vocabulary by comprehensively analyzing the inverse document frequency of knowledge vocabulary and the number of calls in the preset topic. Based on the vocabulary attribution characteristics, knowledge vocabulary is selected and included in the preset topic. Knowledge vocabulary that is not included in the preset topic is screened out and updated in vocabulary classification through the tag inclusion method. The trigger extraction module is used to retrieve the vocabulary coverage and search frequency of knowledge vocabulary when the user searches for knowledge vocabulary to determine whether to trigger key extraction words. When stimulating key extraction words, the trigger value is set to filter the words under the corresponding topic of knowledge vocabulary for provision, and the trigger time is set to calculate the trigger time ratio of each relation extraction word. The relationship correction module is used to call the historical search frequency and collection increment of each relationship extraction word in the historical time and calculate the utilization rate of the relationship extraction word. It comprehensively analyzes the relationship chain characteristics of each relationship extraction word and the current knowledge vocabulary based on the comprehensive triggering time ratio and utilization rate, and provides corrections for the relationship extraction words of the current knowledge vocabulary based on the comprehensive trigger value and relationship chain characteristics.