An AI-based Information Retrieval Method and System

By applying AI-based information retrieval methods in the software help system, the problem that users find it difficult to obtain answers accurately is solved, and more efficient and accurate search results are achieved.

CN119357336BActive Publication Date: 2025-06-13江西博微新技术有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411918441.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-06-13
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

Due to the static document form, traditional software help systems make it difficult for users to obtain answers related to the questions accurately.

Method used

Using AI-based information retrieval method, more accurate and efficient retrieval is achieved by building an index library, vector transformation, vector similarity calculation and interactive behavior feature sorting.

Benefits of technology

It improves the matching degree between search results and user query intentions, improves the search speed and efficiency, and ensures that more comprehensive search results are provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119357336B_ABST
    Figure CN119357336B_ABST
Patent Text Reader

Abstract

The present invention provides an AI-based information retrieval method and system. The method includes: constructing an index library, obtaining initial query data, and then obtaining final query data; splitting the system help document into text segments, and converting the text segments into vector identification data; obtaining retrieval vector data corresponding to the final query data to obtain the vector similarity between it and the vector identification data, and selecting a first target text segment through the vector similarity; selecting a second target text segment from the system help document through the initial query data, sorting and scoring the first target text segment and the second target text segment based on interaction behavior characteristics, and then feeding back the index result. Continuously improve the initial query data through semantic analysis to ensure a high matching degree between search and query; by extracting interaction behavior characteristics, unify the evaluation dimensions of the first target text segment and the second target text segment to ensure that retrieval can be completed quickly, accurately, and strongly associated based on user preferences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to an AI-based information retrieval method and system. Background Art

[0002] Traditional software help systems usually refer to the documents, manuals or built-in help functions provided with the software. Users can obtain answers related to the software by searching for questions within the software help system.

[0003] Although such software help systems are very common, in modern software development practices, some of their limitations have begun to emerge. The most prominent problem is that traditional help systems are often in the form of static documents. After users input question keywords, the traditional help system indexes all relevant content according to the question keywords, which easily leads users to obtain answers with low or even no relevance to the questions. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide an AI-based information retrieval method and system, aiming to solve the technical problem that in the prior art, when users search for answers by inputting question keywords in the software help system, due to the static document form of the help system, it is difficult for users to accurately obtain answers corresponding to the questions.

[0005] To achieve the above purpose, in the first aspect, the embodiments of the present application provide an AI-based information retrieval method, including the following steps:

[0006] Construct an index library corresponding to the system help document, obtain initial question data, and obtain final question data based on the initial question data and the index library;

[0007] Slice the system help document into several text segments, perform vector transformation on the text segments to obtain vector identification data corresponding to the text segments, and combine the vector identification data and the text segments into a database;

[0008] Obtain retrieval vector data corresponding to the final question data, obtain the vector similarity between the retrieval vector data and the vector identification data, and select several first target text segments from several text segments through the vector similarity;

[0009] Select several second target text segments from the system help document through the initial question data, extract the interaction behavior characteristics of the first target text segments and the second target text segments, perform sorting and scoring on the several first target text segments and the several second target text segments based on the interaction behavior characteristics, form a candidate queue through the sorting and scoring, and feedback the index result.

[0010] Compared with the prior art, the beneficial effects of the present invention are as follows: through the preset questions obtained based on the initial question data, the initial question data is continuously improved through semantic analysis to form the final question data, ensuring a high matching degree between the search result and the user's query intention; by converting both the final question data and the system help document into vector data, a faster retrieval speed can be provided, improving the retrieval efficiency; by introducing the full-text retrieval result between the initial question data and the system help document, the problem of reduced retrieval accuracy caused by vector matching is compensated, ensuring that a more comprehensive retrieval result can be provided according to the user's needs; and then by extracting the interaction behavior characteristics, the evaluation dimensions of the first target text segment and the second target text segment are unified, ensuring that the retrieval can be completed quickly, accurately, and strongly associated based on the user's preferences.

[0011] Further, the step of constructing an index library corresponding to the system help document includes:

[0012] Generating a number of subordinate keywords based on the system help document, and setting preset questions according to the subordinate keywords;

[0013] Associating different subordinate keywords to the main keyword according to the meaning of the subordinate keywords, and generating an index library based on the main keyword, the subordinate keywords, and the preset questions.

[0014] Furthermore, the step of obtaining the final question data based on the initial question data and the index library includes:

[0015] Performing text cleaning and word segmentation processing on the initial question data to obtain initial question keywords;

[0016] Comparing the initial question keywords with the main keyword to select a first associated keyword from a number of the main keywords;

[0017] Based on the preset question corresponding to the first associated keyword, obtaining an associated answer from the user, and selecting a number of second associated keywords from a number of the subordinate keywords through the associated answer;

[0018] Combining the first associated keyword and a number of the second associated keywords into the final question data.

[0019] Furthermore, the step of splitting the system help document into a number of text segments includes:

[0020] Dividing the system help document into a number of text blocks with paragraphs as the separation points;

[0021] Obtain the importance of different words in the text block, select key words from several of the words based on the importance, and use the key words as nodes to divide the text block into several text segments.

[0022] Furthermore, the calculation formula for the importance is:

[0023] ,

[0024] where, represents the importance of the i-th word in the j-th document, represents the number of occurrences of the i-th word in the j-th document, represents the total number of words in the j-th document, represents the total number of documents, represents the number of documents containing the i-th word, represents the logarithmic function.

[0025] Furthermore, the step of performing vector transformation on the text segment to obtain vector identification data corresponding to the text segment includes:

[0026] Based on the semantic features of the words in the text segment, divide the text segment into several word groups, and convert the words into sub-codes based on the number of words in the word group;

[0027] Combine several of the sub-codes into an initial vector corresponding to the statement in the text segment, obtain a stage vector corresponding to the text segment according to the initial vector, and convert the stage vector into vector identification data.

[0028] Furthermore, the acquisition formula for the vector identification data is:

[0029] ,

[0030] where, represents the vector identification data, represents the number of sub-codes in the stage vector, represents the number of dimensions, represents the stage vector.

[0031] Furthermore, the calculation formula for the vector similarity is:

[0032] ,

[0033] where, represents the vector similarity, represents the component of the vector identification data in the i-th dimension, represents the component of the retrieval vector data in the i-th dimension, Indicates the number of dimensions.

[0034] Furthermore, the interactive behavior features include retrieval association features, user association features, historical browsing features, and content quality features.

[0035] In a second aspect, an embodiment of the present application provides an AI-based information retrieval system, which is applied to the AI-based information retrieval method described in the first aspect above. The system includes:

[0036] An optimization module, configured to build an index library corresponding to the system help document, obtain initial question data, and obtain final question data based on the initial question data and the index library;

[0037] A conversion module, configured to split the system help document into several text segments, perform vector conversion on the text segments to obtain vector identification data corresponding to the text segments, and combine the vector identification data and the text segments into a database;

[0038] A link module, configured to obtain retrieval vector data corresponding to the final question data, obtain the vector similarity between the retrieval vector data and the vector identification data, and select several first target text segments from several of the text segments through the vector similarity;

[0039] An execution module, configured to select several second target text segments from the system help document through the initial question data, extract the interactive behavior features of the first target text segments and the second target text segments, perform sorting and scoring on the several first target text segments and the several second target text segments based on the interactive behavior features, form a candidate queue through the sorting and scoring, and feedback an index result.

[0040] In a third aspect, an embodiment of the present application provides a computer, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the AI-based information retrieval method described in the first aspect above.

[0041] In a fourth aspect, an embodiment of the present application provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the AI-based information retrieval method described in the first aspect above. Description of the Drawings

[0042] Figure 1 It is a flowchart of the AI-based information retrieval method in the first embodiment of the present invention;

[0043] Figure 2 It is a structural block diagram of the AI-based information retrieval system in the second embodiment of the present invention;

[0044] The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. Specific Embodiments

[0045] For the convenience of understanding the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided so that the disclosure of the present invention is thorough and comprehensive.

[0046] It should be noted that when an element is referred to as being "fixedly provided on" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are only for the purpose of illustration.

[0047] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used herein in the description of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0048] Please refer to Figure 1 , the AI-based information retrieval method provided by the first embodiment of the present invention includes the following steps:

[0049] S10: Construct an index library corresponding to the system help document, obtain initial question data, and obtain final question data based on the initial question data and the index library;

[0050] Specifically, the step S10 includes:

[0051] S110: Generate a number of subordinate keywords based on the system help document, and set a preset question according to the subordinate keywords;

[0052] S120: Associate different subordinate keywords with a main keyword according to the meanings of the subordinate keywords, and generate an index library based on the main keyword, the subordinate keywords and the preset question.

[0053] The subordinate keywords characterize the content features of different sentences or paragraphs in the system help document of the system, and the main keywords are uniformly generated based on the subordinate keywords with the same semantic type. The index library includes a number of the main keywords, and each of the main keywords corresponds to a number of the subordinate keywords. Understandably, each of the main keywords also corresponds to a number of preset questions. Taking the system help document of a financial system as an example, the main keyword can be "expenditure", and the subordinate keywords can be "salary", "water and electricity", etc. It should be noted that the main keywords and the subordinate keywords can be defined by themselves according to the characteristics of the system.

[0054] S130: Perform text cleaning and word segmentation on the initial question data to obtain initial question keywords;

[0055] The text cleaning refers to removing irrelevant characters in the initial question data, and the word segmentation is to decompose the initial question data after text cleaning based on a dictionary to form a number of words.

[0056] S140: Compare the initial question keywords with the main keywords to select a first associated keyword from a number of the main keywords;

[0057] S150: Based on the preset questions corresponding to the first associated keyword, obtain associated answers from the user, and select a number of second associated keywords from a number of the subordinate keywords through the associated answers;

[0058] The preset questions correspond to the subordinate keywords one by one. After determining the first associated keyword, by triggering the corresponding preset questions, the specific retrieval requirements of the user can be judged. In this process, the final selection of the second associated keywords can be completed by means of multiple iterative comparisons.

[0059] S160: Combine the first associated keyword and a number of the second associated keywords into final question data;

[0060] S20: Segment the system help document into a number of text pieces, perform vector conversion on the text pieces to obtain vector identification data corresponding to the text pieces, and combine the vector identification data and the text pieces into a database;

[0061] The step S20 includes:

[0062] S210: Divide the system help document into a number of text blocks with paragraphs as the separation points;

[0063] In some embodiments, after obtaining the text block, subsequent steps can be directly processed. However, since the text block is generated based on paragraphs, although more context information is retained, too much content will greatly increase the retrieval difficulty. Therefore, the text block still needs to be split again.

[0064] S220: Obtain the importance of different words in the text block, select key words from several of the words based on the importance, and use the key words as nodes to split the text block into several text segments;

[0065] The calculation formula for the importance is:

[0066] ,

[0067] where, represents the importance of the i-th word in the j-th document, represents the number of occurrences of the i-th word in the j-th document, represents the total number of words in the j-th document, represents the total number of documents, represents the number of documents containing the i-th word, represents the logarithmic function.

[0068] In this embodiment, the word with the highest importance is selected as the key word, and the statements between two key words are all classified under the previous key word to form the text segment.

[0069] S230: Based on the semantic features of the words in the text segment, divide the text segment into several word groups, and convert the words into sub-codes based on the number of words in the word group;

[0070] Assume that each text segment contains gender features: male, female; sports features: football, basketball, badminton. Then male and female form a word group, and male is converted to (10) and female is converted to (01) respectively. Football, basketball, and badminton form a word group, and football is converted to (100), basketball is converted to (010), and badminton is converted to (001) respectively. It can be understood that in this example, the semantic features are behavior features and sports features.

[0071] S240: Combine several sub-codes into an initial vector corresponding to the statement in the text segment, obtain a stage vector corresponding to the text segment according to the initial vector, and convert the stage vector into vector identification data;

[0072] Specifically, obtain the number of word groups in the text segment. If the number of word groups is 1, directly determine the initial vector as the stage vector; if the number of word groups is greater than 1, perform weighted averaging on the sub-codes at the same positions in different initial vectors to form a stage vector.

[0073] Still taking the above content as an example, assume that the content in the text segment is the results of men's football matches today, the results of men's basketball matches today, and the results of women's badminton matches today. Then the initial vectors corresponding to the sentences in the text segment are: (10100), (10010), (01001), and then perform weighted averaging on the three groups of initial vectors.

[0074] The acquisition formula for the vector identification data is:

[0075] ,

[0076] where, represents the vector identification data, represents the number of sub-codes in the stage vector, represents the number of dimensions, represents the stage vector. It should be noted that due to the different numbers of sub-codes in the stage vectors corresponding to different text segments, through matrix transformation, they can all be converted to the same dimension to facilitate subsequent comparison and retrieval.

[0077] S30: Obtain the retrieval vector data corresponding to the final query data, obtain the vector similarity between the retrieval vector data and the vector identification data, and select several first target text segments from several text segments through the vector similarity;

[0078] The acquisition methods of the retrieval vector data and the vector identification data are the same, and will not be elaborated here. The calculation formula for the vector similarity is:

[0079] ,

[0080] where, represents the vector similarity, represents the component of the vector identification data in the i-th dimension, represents the component of the retrieval vector data in the i-th dimension, represents the number of dimensions.

[0081] After obtaining the vector similarity between the retrieval vector data and each vector identification data, compare the different vector similarities with the similarity threshold, and select the text segments corresponding to the vector similarities greater than the similarity threshold as the first target text segments.

[0082] S40: Select several second target text segments from the system help document according to the initial query data, extract the interaction behavior features of the first target text segments and the second target text segments, rank and score the first target text segments and the second target text segments based on the interaction behavior features, form a candidate queue through the ranking score, and feedback the indexing result;

[0083] Understandably, the second target text segments are the retrieval results obtained by performing a full-text search on the system help document according to the initial query data. The interaction behavior features include retrieval association features, user association features, historical browsing features, and content quality features. The retrieval association features refer to the strength of the association between the text segments and the questions retrieved by the user. The user association features refer to the strength of the association between the text segments and the user's previous retrieval data. The historical browsing features refer to the number of times the text segments have been browsed by all users. The content quality features refer to the high or low content quality of the text segments themselves.

[0084] After obtaining the first target text segments and the second target text segments, by obtaining the interaction behavior features, generating the ranking score corresponding to the first target text segments and the second target text segments based on the feature scores of the interaction behavior features and the feature weights set for different interaction behavior features, then ranking the obtained first target text segments and second target text segments according to the ranking score, forming an indexing result, and determining the priority of the indexing feedback to the user according to the different rankings.

[0085] Through the preset query obtained based on the initial query data, continuously improve the initial query data through semantic analysis to form the final query data, ensuring a high matching degree between the search result and the user's query intention; by converting both the final query data and the system help document into vector data, a faster retrieval speed can be provided, improving the retrieval efficiency; by introducing the full-text retrieval result between the initial query data and the system help document, the problem of reduced retrieval accuracy caused by vector matching is compensated, ensuring that a more comprehensive retrieval result can be provided according to the user's needs; and then by extracting the interaction behavior features, unifying the evaluation dimensions of the first target text segments and the second target text segments, ensuring that the retrieval can be completed quickly, accurately, and strongly associated based on the user's preferences.

[0086] Please refer to Figure 2, the second embodiment of the present invention provides an AI-based information retrieval system, which is applied to the AI-based information retrieval method in the above embodiment. Those that have been described will not be repeated. As used hereinafter, terms such as "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0087] The system includes:

[0088] An optimization module 10, configured to build an index library corresponding to the system help document, obtain initial question data, and obtain final question data based on the initial question data and the index library;

[0089] The optimization module 10 includes:

[0090] A first unit, configured to generate a plurality of subordinate keywords based on the system help document, and set a preset question according to the subordinate keywords;

[0091] A second unit, configured to associate different subordinate keywords with a main keyword according to the meaning of the subordinate keywords, and generate an index library based on the main keyword, the subordinate keywords, and the preset question;

[0092] A third unit, configured to perform text cleaning and word segmentation processing on the initial question data to obtain initial question keywords;

[0093] A fourth unit, configured to compare the initial question keywords with the main keyword to select a first associated keyword from a plurality of the main keywords;

[0094] A fifth unit, configured to obtain an associated answer from the user based on the preset question corresponding to the first associated keyword, and select a plurality of second associated keywords from a plurality of the subordinate keywords through the associated answer;

[0095] A sixth unit, configured to combine the first associated keyword and the plurality of second associated keywords into final question data;

[0096] A conversion module 20, configured to divide the system help document into a plurality of text pieces, perform vector conversion on the text pieces to obtain vector identification data corresponding to the text pieces, and combine the vector identification data and the text pieces into a database;

[0097] The conversion module 20 includes:

[0098] A seventh unit, configured to divide the system help document into a plurality of text blocks with paragraphs as separation points;

[0099] The eighth unit is used to obtain the importance degrees of different words in the text block, select key words from several of the words based on the importance degrees, and take the key words as nodes to divide the text block into several text segments;

[0100] The ninth unit is used to divide the text segments into several word groups based on the semantic features of the words in the text segments, and convert the words into sub-codes based on the number of words in the word groups;

[0101] The tenth unit is used to combine several of the sub-codes into an initial vector corresponding to the sentence in the text segment, obtain a stage vector corresponding to the text segment according to the initial vector, and convert the stage vector into vector identification data;

[0102] The link module 30 is used to obtain retrieval vector data corresponding to the final query data, obtain the vector similarity between the retrieval vector data and the vector identification data, and select several first target text segments from several of the text segments through the vector similarity;

[0103] The execution module 40 is used to select several second target text segments from the system help document through the initial query data, extract the interaction behavior features of the first target text segments and the second target text segments, perform sorting and scoring on several of the first target text segments and several of the second target text segments based on the interaction behavior features, form a candidate queue through the sorting and scoring, and feedback the indexing result.

[0104] The present invention also provides a computer, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the AI-based information retrieval method described in the above technical solution is implemented.

[0105] The present invention also provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the AI-based information retrieval method described in the above technical solution is implemented.

[0106] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0107] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.

Claims

1. An AI-based information retrieval method, characterized in that: The following steps are involved: Constructing an index library corresponding to the system help document, obtaining initial question data, and obtaining final question data based on the initial question data and the index library; The step of constructing an index library corresponding to the system help document comprises: Generate a number of subordinate keywords based on the system help document, and set preset questions according to the subordinate keywords; Associating different subordinate keywords with main keywords according to the meanings of the subordinate keywords, and generating an index library based on the main keywords, the subordinate keywords and the preset questions; The step of obtaining final question data based on the initial question data and the index library comprises: Performing text cleaning and word segmentation processing on the initial question data to obtain initial question keywords; Comparing the initial question keyword with the main keyword to select a first related keyword from a plurality of the main keywords; Based on a preset question corresponding to the first related keyword, obtaining a related answer from the user, and selecting a plurality of second related keywords from the plurality of subordinate keywords based on the related answer; combining the first associated keyword and a plurality of the second associated keywords into final question data; Dividing the system help document into a plurality of text pieces, performing vector conversion on the text pieces to obtain vector identification data corresponding to the text pieces, and combining the vector identification data and the text pieces into a database; Acquire retrieval vector data corresponding to the final question data, acquire vector similarity between the retrieval vector data and the vector identification data, and select a plurality of first target text segments from the plurality of text segments according to the vector similarity; A number of second target text pieces are selected from the system help document using the initial question data, interaction behavior features of the first target text piece and the second target text piece are extracted, the number of first target text pieces and the number of second target text pieces are ranked and scored based on the interaction behavior features, a candidate queue is formed through the ranking and scoring, and an index result is fed back.

2. The AI-based information retrieval method according to claim 1, characterized in that: The step of dividing the system help document into a plurality of text pieces comprises: Dividing the system help document into a plurality of text blocks using paragraphs as delimiters; The importance of different words in the text block is obtained, key words are selected from a plurality of the words based on the importance, and the text block is divided into a plurality of text segments with the key words as nodes.

3. The AI-based information retrieval method according to claim 2, characterized in that: The calculation formula of the importance is: , in, represents the importance of the i-th word in the j-th document, represents the number of occurrences of the i-th word in the j-th document, represents the total number of words in the jth document, Indicates the total number of documents. represents the number of documents containing the i-th word, Represents a logarithmic function.

4. The AI-based information retrieval method according to claim 1, characterized in that: The step of performing vector conversion on the text slice to obtain vector identification data corresponding to the text slice comprises: Based on the semantic features of the words in the text piece, the text piece is divided into a plurality of word groups, and the words are converted into sub-codes based on the number of words in the word groups; A plurality of the sub-codes are combined into an initial vector corresponding to the sentence in the text piece, a phase vector corresponding to the text piece is obtained according to the initial vector, and the phase vector is converted into vector identification data.

5. The AI-based information retrieval method according to claim 4, characterized in that: The formula for obtaining the vector identification data is: , in, Represents vector identification data, represents the number of sub-codes in the phase vector, represents the number of dimensions, Represents the phase vector.

6. The AI-based information retrieval method according to claim 1, characterized in that: The calculation formula of the vector similarity is: , in, represents the vector similarity, Represents the component of the vector identification data in the i-th dimension, Represents the component of the ith dimension of the retrieval vector data. Indicates the number of dimensions.

7. The AI-based information retrieval method according to claim 1, characterized in that: The interactive behavior features include retrieval association features, user association features, historical browsing features and content quality features.

8. An AI-based information retrieval system, applied to the AI-based information retrieval method according to any one of claims 1 to 7, characterized in that: The system comprises: An optimization module, used to construct an index library corresponding to the system help document, obtain initial question data, and obtain final question data based on the initial question data and the index library; The optimization module includes: The first unit is used to generate a number of subordinate keywords based on the system help document, and set preset questions according to the subordinate keywords; The second unit is used to associate different subordinate keywords with main keywords according to the meanings of the subordinate keywords, and generate an index library based on the main keyword, the subordinate keywords and the preset question; The third unit is used to perform text cleaning and word segmentation processing on the initial question data to obtain initial question keywords; A fourth unit is used to compare the initial question keyword with the main keyword to select a first related keyword from a plurality of the main keywords; A fifth unit, configured to obtain a related answer from a user based on a preset question corresponding to the first related keyword, and select a plurality of second related keywords from the plurality of subordinate keywords according to the related answer; A sixth unit, configured to combine the first associated keyword and a plurality of the second associated keywords into final question data; A conversion module, used for dividing the system help document into a plurality of text pieces, performing vector conversion on the text pieces to obtain vector identification data corresponding to the text pieces, and combining the vector identification data and the text pieces into a database; A linking module, used to obtain retrieval vector data corresponding to the final question data, obtain vector similarity between the retrieval vector data and the vector identification data, and select a plurality of first target text segments from the plurality of text segments according to the vector similarity; An execution module is used to select a number of second target text pieces from the system help document through the initial question data, extract the interactive behavior features of the first target text piece and the second target text piece, sort and score the number of first target text pieces and the number of second target text pieces based on the interactive behavior features, form a candidate queue through the sorting and scoring, and feed back the index result.

Citation Information

Patent Citations

  • Query retrieval method and device based on long text and electronic equipment

    CN115630137A

  • Information retrieval method and device, electronic equipment and storage medium

    CN118394916A