Information retrieval and guidance method and apparatus based on big data software system
By splitting, matching, analyzing the correlations and calculating the relevance of user search information, accurate search results are generated, solving the problem of redundant information in existing technologies and achieving more efficient information acquisition and improved user experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING AUGUST MELON TECHNOLOGY CO LTD
- Filing Date
- 2025-09-26
- Publication Date
- 2026-07-30
AI Technical Summary
Existing information retrieval methods lack in-depth analysis of the retrieval questions and effective integration and filtering of retrieval results, resulting in retrieval results often containing a large amount of irrelevant or redundant information, making it difficult to meet users' needs for accurate information acquisition in academic research, business decision-making, and other fields.
By splitting, matching, analyzing the relevance of, and comprehensively analyzing user search information, accurate search information is generated, including splitting matching results, integrating related content, analyzing duplicate information, and calculating relevance, ultimately generating ordered search results.
It improves the targeting and accuracy of searches, reduces redundant information, broadens the scope of information access, provides users with a more comprehensive and in-depth knowledge system, and improves the user experience.
Smart Images

Figure CN2025124268_30072026_PF_FP_ABST
Abstract
Description
A method and apparatus for information retrieval and guidance based on a big data software system This application claims priority to Chinese Patent Application No. 202510114548.7, filed on January 24, 2025, entitled "A Method and Apparatus for Information Retrieval and Guidance Based on a Big Data Software System", the entire contents of which are incorporated herein by reference. Technical Field
[0001] This invention relates to the field of data processing technology, specifically to a method and apparatus for information retrieval and guidance based on a big data software system. Background Technology
[0002] In the field of information retrieval, with the explosive growth of data volume, traditional information retrieval methods face numerous challenges. When faced with massive amounts of information, users often find it difficult to quickly and accurately obtain comprehensive information that is highly relevant to their search questions.
[0003] Patent application CN115510085A discloses a method and apparatus for information retrieval and guidance based on a big data software system, relating to the field of data processing technology. The method includes: acquiring user-input keywords; processing the keywords using a preset tag attribute model to obtain keyword retrieval results; wherein the retrieval results include: precisely matched functional field information or fuzzy matched functional field information; the functional field information includes: the functional field's tag, the functional field's name, the functional field's meaning, and the functional field's Uniform Resource Locator (URL); responding to the user's access to the URL, redirecting to the management page corresponding to the functional field, and guiding the user to use the management functions corresponding to the functional field in the big data software system based on the management page.
[0004] However, most existing search systems are based on simple keyword matching, lacking in-depth analysis of the search questions and effective integration and filtering of search results. This results in search results often containing a large amount of irrelevant or redundant information, requiring users to spend a lot of time and effort to sift through the complex results to find useful content. Moreover, for complex search needs, they cannot effectively explore the relationships between different parts, making it difficult to provide systematic and targeted information, and failing to adequately meet users' growing demand for accurate information in academic research, business decision-making, and professional technical exploration. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a method and apparatus for information retrieval and guidance based on a big data software system. This solves the problem that the lack of in-depth analysis of the retrieval problem and effective integration and filtering of retrieval results often leads to retrieval results containing a large amount of irrelevant or redundant information.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for information retrieval and guidance based on a big data software system, comprising the following steps: Step 1: Acquiring the user's retrieval information, and simultaneously decomposing the retrieval question into decomposed retrieval content based on the retrieval information; matching the decomposed retrieval content with the database to obtain decomposed matching results; and filtering out identical matching results based on the decomposed matching results; Step 2: Obtaining the remaining decomposed matching results, and simultaneously performing correlation analysis to obtain related content; and integrating the related content to obtain the results to be analyzed; Step 3: Performing comprehensive analysis on the obtained identical matching results and the results to be analyzed to obtain comprehensive analysis results; integrating the comprehensive analysis results corresponding to the same semantics to obtain secondary classification results; and simultaneously analyzing the duplicate information corresponding to the secondary classification results to generate duplicate analysis results; Step 4: Analyzing the obtained duplicate analysis results; filtering by calculating the relevance between the duplicate analysis results and the retrieval question; and sorting them from largest to smallest according to the calculated relevance to generate retrieval information.
[0007] As a further aspect of the present invention, the specific method for obtaining the same matching result in step one is as follows: The search information of the user is obtained, and then the search question is split according to the topic content to obtain split search content. The split search content is obtained and labeled as n, where n = 1, 2, ..., m, and m represents the number of split search content. Then, the split search content n is matched with the corresponding database to obtain the corresponding split matching results. The split matching results corresponding to different split search content are obtained, and then the split matching results are categorized to obtain the same matching result.
[0008] As a further aspect of the present invention, the specific method for obtaining the analysis result in step two is as follows: All remaining split matching results are obtained, and the remaining split matching results are distinguished and marked as split matching result a, where a = A, B, ..., where a represents the type of split search content corresponding to the remaining split matching result. Then, any type of split matching result is obtained as the analysis object, and the corresponding matching result in the analysis object is denoted as i, where i = 1, 2, ..., j, where j represents the number of matching results. Simultaneously, the matching content of matching result i is identified, and then matching result i is performed with the database to obtain associated content. This process is repeated for all split matching results to obtain the corresponding associated content. Next, similarity analysis is performed on the associated content corresponding to different split matching results, and associated content with similar content is filtered to generate the analysis result.
[0009] As a further embodiment of the present invention, the specific method for obtaining the secondary classification result in step three is as follows: obtain all identical matching results and results to be analyzed, and combine the two to obtain a comprehensive analysis result, which is labeled as c, and c = 1, 2, ..., r, where r represents the number of comprehensive analysis results. Then, the semantic content corresponding to the comprehensive analysis result c is analyzed, and the semantic similarity of the content is calculated. The comprehensive analysis results with the same semantics are integrated to obtain the secondary classification result. At the same time, the number of results corresponding to the secondary classification result is obtained, and the proportion of the number corresponding to the secondary classification result is calculated. Then, the results are sorted from largest to smallest according to the proportion of the number.
[0010] As a further aspect of the present invention, the specific method for obtaining the duplicate analysis results in step three is as follows: Obtain all secondary classification results, and simultaneously obtain any group of secondary classification results as the target object, and obtain the corresponding comprehensive analysis results in the target object. Then, extract duplicate information from the comprehensive analysis results, and obtain the number of duplicate information corresponding to the comprehensive analysis results in the target object. Use the duplicate information as a standard to obtain the corresponding comprehensive analysis results and record them as identical results. Repeat this process, analyzing all duplicate information and generating corresponding identical results. Next, obtain the types of duplicate information for identical results. If there is only one type of duplicate information corresponding to the identical result, then remove the identical result. If there are multiple types of duplicate information corresponding to the identical result, obtain and retain the comprehensive analysis results corresponding to multiple types of duplicate information, and simultaneously remove the remaining comprehensive analysis results in the identical results to generate duplicate analysis results.
[0011] As a further aspect of the present invention, the specific method for generating retrieval information in step four is as follows: All duplicate analysis results are obtained and labeled as k, where k = 1, 2, ..., h, and h represents the number of duplicate analysis results. Then, the retrieval question is obtained, and the identical words present in the duplicate analysis results are obtained using the retrieval question as the standard. The similarity value between the duplicate analysis results and the retrieval question is also obtained. Then, the number of identical words and the similarity value are substituted into the formula Q = (S + G) × u to calculate the relevance of the duplicate analysis results, where u is a preset proportional coefficient, S is the number of identical words, and G is the similarity value. Similarly, the relevance Qk corresponding to all duplicate analysis results is calculated, and the mean of all relevance values is calculated. Simultaneously, the relevance Qk is used as the standard for filtering, removing duplicate analysis results with a relevance Qk less than the mean relevance value, and retaining duplicate analysis results with a relevance Qk greater than the mean relevance value. Finally, the relevance Qk is sorted from largest to smallest to generate retrieval information.
[0012] An information retrieval and guidance device based on a big data software system includes an information acquisition module, an information matching and analysis module, a comprehensive analysis and processing module, and a retrieval information output module. The information acquisition module acquires user retrieval information and transmits it to the information matching and analysis module. The information matching and analysis module breaks down the retrieval question into split retrieval content based on the retrieval information, matches the split retrieval content with a database to obtain split matching results, filters identical matching results based on the split matching results, performs correlation analysis to obtain related content, integrates the related content to obtain the result to be analyzed, and then transmits the result to the comprehensive analysis and processing module. The comprehensive analysis and processing module performs comprehensive analysis on the obtained identical matching results and the result to be analyzed to obtain a comprehensive analysis result. It integrates the comprehensive analysis results corresponding to the same semantics to obtain a secondary classification result, analyzes the duplicate information corresponding to the secondary classification result to generate duplicate analysis results, filters the duplicate analysis results by calculating the relevance between the content and the retrieval question, sorts them according to the calculated relevance from largest to smallest, generates retrieval information, and transmits the retrieval information to the retrieval information output module. The retrieval information output module is used to display the retrieved retrieval information to the corresponding operators.
[0013] This invention provides a method and apparatus for information retrieval and guidance based on a big data software system. Compared with existing technologies, it has the following advantages: By finely breaking down the retrieval question according to the topic content, this invention can more accurately understand the user's retrieval intent. Compared with the traditional retrieval method based on overall keywords, it improves the targeting and accuracy of the retrieval, avoids the confusion and inaccuracy of retrieval results caused by the general use of keywords, and the semantic analysis method can more accurately identify information with similar content, overcome the limitations of traditional word matching, effectively reduce redundant information, and make the retrieval results more concise and valuable. By analyzing related content to mine potential related content and integrating it into the results to be analyzed, it can reintegrate and utilize relevant information that might otherwise be ignored, broaden the scope of information access for users, and provide users with a more comprehensive and in-depth knowledge system.
[0014] By calculating the relevance between the duplicate analysis results and the search question, and then filtering and ranking them based on the relevance, a scientific quantitative method is used to ensure that the search information presented to users is highly relevant to the original search question. This improves the efficiency of users obtaining useful information and enhances the user experience. Traditional technologies often struggle to effectively quantify and rank the relevance between search results and the question. Attached Figure Description
[0015] Figure 1 is a flowchart of the steps of the present invention; Figure 2 is a block diagram of the system principle of the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] Example 1, please refer to Figure 1. This application provides an information retrieval and guidance method based on a big data software system. The method specifically includes the following steps: Step 1: Obtain the user's retrieval information, and at the same time, split the retrieval question according to the retrieval information to obtain split retrieval content, match the obtained split retrieval content with the database to obtain split matching results, and at the same time, filter the same matching results according to the split matching results.
[0018] First, the search information of the user is obtained, including the user's search question. Then, the search question is broken down according to the topic content to obtain the split search content.
[0019] For example, if a user enters the search query "research materials on the current status and development trend of artificial intelligence in the medical field", it can be further broken down into multiple sub-search contents such as "current status of artificial intelligence in the medical field", "development trend of artificial intelligence in the medical field" and "research materials on the medical application of artificial intelligence" based on the inherent logic and semantic relationship of the topic content.
[0020] The split search content is obtained and labeled as n, where n = 1, 2, ..., m, and m represents the number of split search content. Then, each split search content n is matched with the corresponding database to obtain the corresponding split matching results. The split matching results obtained here are the matching results corresponding to different split search content individually. The split matching results corresponding to different split search content are obtained. Then, the split matching results are categorized to obtain the same matching results. Specifically, the categorization process here is to classify the split matching results with similar content. The specific content similarity is calculated by calculating the content similarity value using the cosine similarity calculation method.
[0021] Step 2: Obtain the remaining split matching results, perform correlation analysis to obtain related content, and integrate the related content to obtain the results to be analyzed.
[0022] Retrieve all remaining split matching results, where "remaining split matching results" refers to the split matching results remaining after removing identical matching results. Specifically, this includes different split matching results corresponding to the split search content. For example, if there are two types of split search content, the corresponding different split matching results are the matching results corresponding to the two types of split search content. The remaining split matching results are distinguished and marked as split matching result 'a', where a = A, B, ... The meaning of the matching result here is the same as that of the remaining split matching results, where 'a' represents the type of split search content corresponding to the remaining split matching results. Next, obtain any type of split matching result as the analysis object, and denote the corresponding matching result in the analysis object as i, where i = 1, 2, ..., j, where j represents the number of matching results. At the same time, identify the matching content of matching result i, and then perform correlation matching between matching result i and the database. Here, correlation matching means performing correlation matching on the corresponding content to obtain related content. This process is repeated for all split matching results to obtain the corresponding related content. Then, perform similarity analysis on the related content corresponding to different split matching results, and filter out related content with similar content to generate the analysis results.
[0023] Step 3: Perform a comprehensive analysis on the obtained identical matching results and the results to be analyzed to obtain a comprehensive analysis result. Integrate the comprehensive analysis results corresponding to the same semantics to obtain a secondary classification result. At the same time, analyze the duplicate information corresponding to the secondary classification results to generate a duplicate analysis result.
[0024] All identical matching results and results to be analyzed are obtained and combined to obtain a comprehensive analysis result, which is labeled as c, and c = 1, 2, ..., r, where r represents the number of comprehensive analysis results. Then, the semantic content corresponding to the comprehensive analysis result c is analyzed, and the semantic similarity is calculated. The comprehensive analysis results with the same semantics are integrated to obtain the secondary classification result. The secondary classification result here specifically represents the comprehensive analysis results with the same semantic content. At the same time, the number of results corresponding to the secondary classification result is obtained, and the proportion of the number corresponding to the secondary classification result is calculated. Then, they are sorted from largest to smallest according to the proportion of the number. The process involves obtaining all secondary classification results, taking any set of secondary classification results as the target object, and obtaining the corresponding comprehensive analysis results for the target object. Next, duplicate information is extracted from the comprehensive analysis results. Duplicate information here refers to identical information present in different comprehensive analysis results; specifically, this could be identical information present in two or three sets of comprehensive analysis results. The number of duplicate information types corresponding to the comprehensive analysis results in the target object is also obtained. The corresponding comprehensive analysis results based on duplicate information are recorded as identical results. This process continues, analyzing all duplicate information and generating corresponding identical results. Then, the types of duplicate information in the identical results are obtained. If only one type of duplicate information exists for an identical result, the identical result is removed. If multiple types of duplicate information exist for an identical result, comprehensive analysis results with multiple types of duplicate information are obtained and retained. The remaining comprehensive analysis results in the identical results are then removed, generating duplicate analysis results.
[0025] Specifically, for comprehensive analysis results containing multiple duplicate information, such as five groups, each with different types of duplicate information, the result with the most duplicate information is selected as the standard for retention, while the rest are directly discarded.
[0026] Step 4: Analyze the obtained duplicate analysis results, filter them by calculating the relevance between the duplicate analysis results and the search question, sort them from largest to smallest according to the calculated relevance, and generate search information.
[0027] All duplicate analysis results are retrieved and labeled k, where k = 1, 2, ..., h, and h represents the number of duplicate analysis results. Next, the search question is retrieved, and identical words in the duplicate analysis results are identified using the search question as the criterion. Here, identical words refer to word categories. The similarity value between the duplicate analysis results and the search question is calculated using the Euclidean distance formula. The number of identical words and the similarity value are then substituted into the formula Q = (S + G) × u to calculate the relevance of the duplicate analysis results, where u is a preset proportional coefficient, S is the number of identical words, and G is the similarity value. This process is repeated for all duplicate analysis results, calculating the relevance Qk. The mean of all relevance values is then calculated, and the results are filtered based on the relevance Qk. Duplicate analysis results with a relevance Qk less than the mean are removed, while those with a relevance Qk greater than the mean are retained. Finally, the results are sorted from largest to smallest relevance Qk, generating search information, which is then displayed to the relevant operators.
[0028] Example 2, please refer to Figure 2. This application provides an information retrieval and guidance device based on a big data software system. The device specifically includes an information acquisition module, an information matching and analysis module, a comprehensive analysis and processing module, and a retrieval information output module.
[0029] The information acquisition module acquires the user's search information and transmits it to the information matching and analysis module. The information matching and analysis module breaks down the search question based on the search information to obtain split search content. It then matches the split search content with the database to obtain split matching results, filters out identical matching results based on the split matching results, performs correlation analysis to obtain related content, integrates the related content to obtain the results to be analyzed, and then transmits the results to the comprehensive analysis and processing module. The processing method here is the same as steps two and three in Embodiment 1. The comprehensive analysis and processing module performs comprehensive analysis on the obtained identical matching results and the results to be analyzed to obtain comprehensive analysis results. It integrates the comprehensive analysis results corresponding to the same semantics to obtain secondary classification results, analyzes the duplicate information corresponding to the secondary classification results to generate duplicate analysis results, filters the duplicate analysis results by calculating the relevance between the content and the search question, sorts them according to the calculated relevance from largest to smallest, generates search information, and transmits the search information to the search information output module. The processing method here is the same as steps three and four in Embodiment 1. The retrieval information output module is used to display the retrieved retrieval information to the corresponding operators.
[0030] Some of the data in the above formulas are numerical calculations with dimensions removed, and the contents not described in detail in this specification are all prior art known to those skilled in the art.
[0031] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.
Claims
1. A method for information retrieval and guidance based on big data software system, characterized in that, The method comprises the following steps: Step one: the retrieval information of the user is acquired, and the retrieval question is split according to the retrieval information to obtain split retrieval content; the split retrieval content is matched with a database to obtain split matching results; and the same matching results are obtained by screening the split matching results; Step two: the remaining split matching results are acquired, the associated content is obtained by analyzing the correlation of the content, and the associated content is integrated to obtain the results to be analyzed; Step three: the same matching results and the results to be analyzed are comprehensively analyzed to obtain comprehensive analysis results; the comprehensive analysis results corresponding to the same semantics are integrated to obtain secondary classification results; and the repeated information corresponding to the secondary classification results is analyzed to generate repeated analysis results; Step four: the repeated analysis results are analyzed, the relevance of the repeated analysis results and the retrieval question is calculated, the repeated analysis results are screened according to the relevance, and the retrieval information is generated by sorting the relevance from large to small.
2. The method of information retrieval and guidance based on big data software system according to claim 1, wherein, The specific way of obtaining the same matching results in step one is as follows: The retrieval information of the retrieval user is acquired, the retrieval question is split according to the theme content to obtain split retrieval content, the split retrieval content is labeled as n, and n=1, 2, …, m, wherein m represents the number of split retrieval content; the split retrieval content n is matched with the corresponding database, and the corresponding split matching results are obtained; the split matching results corresponding to different split retrieval content are obtained; and the split matching results are classified to obtain the same matching results.
3. The method of information retrieval and guidance based on big data software system according to claim 1, wherein, The specific way of obtaining the results to be analyzed in step two is as follows: All the remaining split matching results are acquired, and the remaining split matching results are distinguished and labeled as split matching results a, and a=A, B, …, wherein a represents the type of split retrieval content corresponding to the remaining split matching results; any type of split matching result is taken as an analysis object, and the matching result corresponding to the analysis object is labeled as i, and i=1, 2, …, j, wherein j represents the number of matching results; the matching content of the matching result i is identified; the matching result i is associated with the database to obtain the associated content; and all the split matching results are analyzed to obtain the corresponding associated content. The specific way of obtaining the secondary classification results in step three is as follows:
4. The method of information retrieval and guidance based on big data software system according to claim 1, wherein, All the same matching results and the results to be analyzed are acquired, and the two are combined to obtain comprehensive analysis results, which are labeled as c, and c=1, 2, …, r, wherein r represents the number of comprehensive analysis results; the content semantics corresponding to the comprehensive analysis results c are analyzed, the similarity of the content semantics is calculated, the comprehensive analysis results with the same semantics are integrated to obtain the secondary classification results, the number of the secondary classification results is acquired, the proportion of the number of the secondary classification results is calculated, and the secondary classification results are sorted according to the proportion from large to small. 5. The method of information retrieval and guidance based on big data software system according to claim 1, wherein, The specific way of obtaining the repeated analysis result in step three is as follows: All secondary classification results are obtained, and any group of secondary classification results is taken as a target object, and the corresponding comprehensive analysis result in the target object is obtained. Then, the repeated information in the comprehensive analysis result is extracted, the number of repeated information corresponding to the comprehensive analysis result in the target object is obtained, and the corresponding comprehensive analysis result is recorded as the same result as the repeated information. In this way, all repeated information is analyzed, and the corresponding same result is generated. Then, the types of repeated information of the same result are obtained. If there is only one type of repeated information corresponding to the same result, the same result is removed. If there are multiple types of repeated information corresponding to the same result, the comprehensive analysis results corresponding to the multiple types of repeated information are obtained, and are retained at the same time. The remaining comprehensive analysis results in the same result are removed to generate the repeated analysis result.
6. The method of information retrieval and guidance based on big data software system according to claim 1, wherein, The specific way of generating the search information in step four is as follows: All repeated analysis results are obtained and labeled as k, and k=1, 2, …, h, where h represents the number of repeated analysis results. Then, the search question is obtained, and the same words in the repeated analysis result are obtained according to the search question. The similarity value between the repeated analysis result and the search question is obtained. Then, the number of same words and the similarity value are substituted into the formula Q=(S+G)×u to calculate the relevance of the repeated analysis result, where u is a preset proportion coefficient, S is the number of same words, and G is the similarity value. In this way, the relevance Qk of all repeated analysis results is calculated, and the average value of all relevance sums is calculated. According to the average value, the repeated analysis results with a relevance Qk less than the average value of the relevance are removed, and the repeated analysis results with a relevance Qk greater than the average value of the relevance are retained. Then, the repeated analysis results are sorted in descending order of the relevance Qk to generate the search information.
7. An apparatus for information retrieval and guidance based on a big data software system, configured to perform a method of information retrieval and guidance based on a big data software system according to any one of claims 1 to 6, characterized in that, The information collection module, the information matching analysis module, the comprehensive analysis processing module, and the search information output module are included. 8.The information retrieval and guidance device based on big data software system of claim 7, wherein, The information collection module is used to obtain the search information of the user and transmit the search information to the information matching analysis module. The information matching analysis module is used to split the search question according to the search information to obtain split search content, match the split search content with the database to obtain split matching results, filter the same matching results according to the split matching results, analyze the association content to obtain associated content, integrate the associated content to obtain analysis results, and transmit the analysis results to the comprehensive analysis processing module. The comprehensive analysis processing module performs comprehensive analysis on the obtained same matching results and the to-be-analyzed results to obtain comprehensive analysis results, integrates the comprehensive analysis results corresponding to the same semantics to obtain secondary classification results, analyzes repeated information corresponding to the secondary classification results to generate repeated analysis results, filters the repeated analysis results by calculating the relevance of the contents of the repeated analysis results to the search question, sorts the repeated analysis results in descending order according to the calculated relevance, generates search information, and transmits the search information to the search information output module; The search information output module is configured to display the obtained search information to corresponding operating personnel.