An accurate search method and system for massive Internet data based on AI technology

By generating structured query vectors, building knowledge graphs and personalized search models, the shortcomings of traditional search engines in semantic understanding and personalized recommendations are solved, and a deep understanding of user intentions and personalized search is achieved, which improves the efficiency and user experience of information retrieval.

CN119557500BActive Publication Date: 2025-07-18BEIJING DINGXI YINGDONG TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411518055.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-29
Publication Date
2025-07-18
Estimated Expiration
2044-10-29

AI Technical Summary

Technical Problem

When traditional search engines process complex queries and diversified information, their semantic understanding capabilities are limited, and they lack personalized recommendations and dynamic adjustments to users' real-time behavior, resulting in inefficient and overloaded cognitive load when users acquire information from massive data.

Method used

Natural language processing technology is used to generate structured query vectors, build a knowledge graph, introduce knowledge transfer learning technology to analyze user behavior, monitor cognitive load and generate adaptive search interfaces, and initiate multiple query requests using a personalized search model.

Benefits of technology

It has achieved a deep understanding of user query intentions, provided personalized, relevant and diverse search results, reduced user cognitive load, and improved information retrieval efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119557500B_ABST
    Figure CN119557500B_ABST
Patent Text Reader

Abstract

The present invention discloses an accurate search method and system for massive Internet data based on AI technology, which relates to the field of Internet technology. It includes receiving a query text input by a user, and using natural language processing technology to identify the input query text to generate a structured query vector; based on the structured query vector, constructing a knowledge graph to obtain a rich semantic context graph; introducing knowledge transfer learning technology to analyze the user's historical search data and behavioral characteristics to construct a personalized search model; monitoring the user's real-time interaction behavior, judging the user's cognitive load, and generating an adaptive search interface; using the personalized search model to initiate multiple query requests, and after filtering and sorting, outputting a personalized search result set. By constructing a knowledge graph and generating a rich semantic context graph, the system of the present invention can identify entities and their multiple relationships, thereby providing more relevant search results and promoting users to discover potential needs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet technologies, and in particular to an accurate search method and system for Internet massive data based on AI technology. Background Art

[0002] With the rapid development of information technology, the Internet has become a huge data resource library, and the generation and storage of massive data have put forward higher requirements for information retrieval technology. Traditional search engines mainly rely on keyword matching and static algorithms to process user queries. In the face of complex queries and diverse information, this method often seems inadequate. In recent years, the rise of emerging technologies such as natural language processing (NLP), machine learning, and knowledge graphs has provided new ideas for accurate search.

[0003] Although the existing technologies have made certain progress in the field of information retrieval, there are still many deficiencies. First, the traditional search engine has limited ability to understand the semantics of queries. Especially when dealing with fuzzy or long-tail queries, it is often difficult to provide highly relevant results. This results in users often having to conduct multiple queries during the information acquisition process, consuming time and energy. Second, the existing personalized search technologies mostly rely on static models and lack the ability to monitor and dynamically adjust the real-time behavior of users, thus making it difficult to adapt to the changes in user needs. In addition, the cognitive load of users is often ignored during the information retrieval process, resulting in users feeling confused when facing massive information, further affecting the search efficiency and satisfaction. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides an accurate search method and system for Internet massive data based on AI technology to solve the problems of insufficient query understanding, lack of personalized recommendation, and excessive cognitive load faced by users when obtaining accurate information from massive data.

[0006] To solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, an embodiment of the present invention provides an accurate search method for Internet massive data based on AI technology, which includes receiving a query text input by a user and using natural language processing technology to identify the input query text to generate a structured query vector;

[0008] Based on the structured query vector, construct a knowledge graph to obtain a rich semantic context graph;

[0009] Introduce knowledge transfer learning technology, analyze the historical search data and behavior characteristics of users, and construct a personalized search model;

[0010] Monitor the real-time interaction behavior of users, judge the cognitive load of users, and generate an adaptive search interface;

[0011] Initiate multiple query requests using a personalized search model, and after filtering and sorting, output a personalized search result set.

[0012] As a preferred solution of the Internet mass data precise search method based on AI technology described in the present invention, wherein: receiving the query text input by the user, and using natural language processing technology to identify the input query text, generating a structured query vector includes the following steps,

[0013] The user inputs a query text in the search interface;

[0014] Use a dependency syntax analysis tool to perform basic syntax analysis on the input text;

[0015] Construct a prefix dictionary of common words, and perform word segmentation on the analyzed text;

[0016] Perform part-of-speech tagging on the text after word segmentation;

[0017] After completing part-of-speech tagging, perform named entity recognition on the text;

[0018] Combine the results of word segmentation, part-of-speech tagging, and named entity recognition to perform keyword extraction and topic recognition;

[0019] Combine the extracted keywords, user intentions, and their related attributes into a structured query vector and store it in the knowledge base.

[0020] As a preferred solution of the Internet mass data precise search method based on AI technology described in the present invention, wherein: the basic syntax analysis refers to using a dependency syntax analysis tool to parse the input text, identify each word in the sentence and its corresponding part of speech, automatically analyze the relationship between words, and generate a dependency relationship graph; the word segmentation processing refers to using a maximum word length word segmentation algorithm based on a prefix dictionary to split the input string into independent words to obtain a word segmentation result;

[0021] The specific steps of the part-of-speech tagging are as follows,

[0022] Select conditional random field as the model for part-of-speech tagging;

[0023] Extract the current word features, context features, and part-of-speech features from the word segmentation results;

[0024] Pass each word segmentation and its features to the conditional random field model, calculate the part-of-speech probability of each word, and the expression is:

[0025]

[0026] Among them, P(y i | w i ) represents the probability of the part-of-speech tag y i given the word segmentation w i , where w i represents the i-th word segmentation, y i represents the part-of-speech tag associated with the i-th word segmentation w i , i represents the word segmentation index, θ represents the parameter vector of the model, f represents the set of features, T represents the transpose operation, Y represents the set of all possible part-of-speech tags, and y' represents the tag in the set of part-of-speech tags Y;

[0027] For each word segmentation, select the one with the highest probability from all possible part-of-speech tags as the part-of-speech annotation of this word segmentation;

[0028] The process of the named entity recognition is as follows.

[0029] Select a model that combines a bidirectional long short-term memory network and a conditional random field for named entity recognition;

[0030] Input the word segmentations with part-of-speech annotations into the model for training;

[0031] Pass the word segmentation results and part-of-speech annotations to the trained model, perform named entity recognition on each word, and output the corresponding entity category;

[0032] The specific steps of the keyword extraction and topic recognition are as follows.

[0033] Adopt the TF-IDF algorithm to calculate the TF-IDF value of each word, and the expression is

[0034] TF-IDF(t, d) = TF(t, d) × IDF(t);

[0035] Among them, TF-IDF(t, d) represents the TF-IDF value of the word t in the document d, t represents a specific word, d represents a specific document, TF(t, d) represents the word frequency of the word t in the document d, and IDF(t) represents the inverse document frequency of the word t;

[0036] Sort all words according to the TF-IDF value, and select the top N words with the highest TF-IDF value as keywords;

[0037] Identify the main topic of the text by analyzing the relationships and contexts among the keywords.

[0038] As a preferred solution of the Internet mass data precise search method based on AI technology described in the present invention, wherein: based on the structured query vector, constructing a knowledge graph to obtain a rich semantic context graph includes the following steps,

[0039] Connect to the selected knowledge base through a database connection tool;

[0040] Extract relevant metadata according to the user's query conditions;

[0041] Based on the extracted keywords and themes, perform keyword matching in the knowledge base to identify multiple entities related to the user's query;

[0042] According to the identified multiple relevant entities, retrieve the relationships between these entities;

[0043] Classify the obtained relationships and extract relevant attributes to judge the interaction between entities;

[0044] Create a node for each identified entity and create the identified relationships between entities;

[0045] Construct the identified entities and their relationships into a knowledge graph;

[0046] For each node in the knowledge graph, initialize its feature vector;

[0047] Introduce the message passing mechanism of the graph neural network, and the node passes its feature vector to adjacent nodes;

[0048] Set the number of inference layers K, and through multi-layer graph neural network inference, gradually update the feature vector of each node by combining the feature vectors of its neighbor nodes to obtain the final feature vector of each node;

[0049] Analyze the final feature vector of each node, identify other topics and concepts similar to the user's query intention, and construct a rich semantic context graph.

[0050] As a preferred solution of the Internet mass data precise search method based on AI technology described in the present invention, wherein: introducing knowledge transfer learning technology, analyzing the user's historical search data and behavior characteristics, and constructing a personalized search model includes the following steps,

[0051] Extract the feature vector of each node and the feature vector of the edge from the semantic context graph;

[0052] Integrate the extracted node feature vector and edge feature vector into an input matrix;

[0053] Collect the user's historical search data and behavior characteristics and perform feature extraction to form a user feature matrix;

[0054] Concatenate the input feature matrix and the user feature matrix to form a comprehensive feature matrix;

[0055] Construct an adjacency matrix to represent the connection relationship between nodes in the graph;

[0056] Select a transfer learning model suitable for processing graph-structured data, input the comprehensive feature matrix and the adjacency matrix into the transfer learning model to fine-tune the model, and construct a personalized search model.

[0057] As a preferred solution of the Internet massive data precise search method based on AI technology described in the present invention, wherein: monitoring the real-time interaction behavior of the user, judging the cognitive load of the user, and generating an adaptive search interface includes the following steps,

[0058] Collect the interaction behavior data of the user on the personalized search result page in real time;

[0059] Evaluate the cognitive load of the user according to the interaction behavior data of the user;

[0060] Formulate a dynamic interface adjustment strategy based on the evaluation result of the user's cognitive load;

[0061] Adjust the layout of the interface for different cognitive load levels to generate an adaptive search interface.

[0062] As a preferred solution of the Internet massive data precise search method based on AI technology described in the present invention, wherein: using the personalized search model to initiate multiple query requests, and after filtering and sorting, outputting a personalized search result set includes the following steps,

[0063] Use the personalized search model to analyze the user's query intention and generate multiple query requests;

[0064] Send the generated query requests to the knowledge base and related databases to capture information related to the query requests;

[0065] According to the user's personalized characteristics and the evaluation result of the cognitive load, preliminarily filter the captured results to remove irrelevant information;

[0066] Sort the filtered results, evaluate the diversity of the results, and generate a final search result set.

[0067] In a second aspect, the present invention provides an Internet massive data precise search system based on AI technology, including a query input module, which receives the query text input by the user and uses natural language processing technology to identify the input query text to generate a structured query vector;

[0068] A knowledge graph construction module constructs a knowledge graph based on the structured query vector to obtain a rich semantic context graph;

[0069] The knowledge transfer learning module introduces knowledge transfer learning technology to analyze users' historical search data and behavior characteristics and build a personalized search model;

[0070] The cognitive load assessment module monitors the user's real-time interactive behavior, determines the user's cognitive load, and generates an adaptive search interface;

[0071] The query request generation module uses the personalized search model to initiate multiple query requests, and outputs a personalized search result set after filtering and sorting.

[0072] In a third aspect, an embodiment of the present invention provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the method for precise search of massive Internet data based on AI technology as described in the first aspect of the present invention.

[0073] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the method for precise search of massive Internet data based on AI technology as described in the first aspect of the present invention.

[0074] The beneficial effects of the present invention are as follows: by receiving the query text input by the user and generating a structured query vector using natural language processing technology, a deep understanding of the user's query intention is achieved; a knowledge graph is constructed and a rich semantic context graph is generated, so that the system can identify entities and their multiple relationships, thereby providing more relevant search results and promoting users to discover potential needs; knowledge transfer learning technology is introduced, and a personalized search model is constructed by analyzing the user's historical search data, so that accurate recommendations for individual users are achieved, the relevance of the search is improved, and the user's cognitive load is reduced; by real-time monitoring of user interaction behavior, an adaptive search interface is generated, and the user experience is dynamically optimized, making information retrieval more efficient and friendly; a personalized search model is used to initiate multiple query requests, and after filtering and sorting, a relevant and diverse search result set is output to ensure that users can efficiently obtain information. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0076] Figure 1 It is a flowchart of the precise search method for Internet massive data based on AI technology in Embodiment 1.

[0077] Figure 2 It is a flowchart of constructing a personalized search model in Embodiment 1. Specific implementation manners

[0078] To make the above objects, features, and advantages of the present invention more obvious and understandable, the specific implementation manners of the present invention will be described in detail below with reference to the accompanying drawings of the specification.

[0079] In the following description, many specific details are set forth to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0080] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it an embodiment that is separate or selectively exclusive of other embodiments.

[0081] Embodiment 1, referring to Figure 1 and Figure 2 , is the first embodiment of the present invention. This embodiment provides a precise search method for Internet massive data based on AI technology, including the following steps:

[0082] S1. Receive the query text input by the user, and use natural language processing technology to identify the input query text, and generate a structured query vector, including the following steps.

[0083] S1.1. The user inputs the query text on the search interface; use a dependency syntax analysis tool to perform basic grammar analysis on the input text.

[0084] Specifically, the basic grammar analysis refers to using a dependency syntax analysis tool to parse the input text, identify each word in the sentence and its corresponding part of speech, automatically analyze the relationship between the words, and generate a dependency relationship graph.

[0085] It should be noted that through dependency syntax analysis, the grammatical relationship between words in the text can be clearly identified, laying a foundation for subsequent semantic understanding. The text analysis methods mentioned in the prior art mainly rely on basic spelling checks and simple word segmentations, and fail to deeply analyze the sentence structure, which may lead to information loss and understanding deviation. Dependency syntax analysis can provide a more detailed understanding of the sentence structure, thereby improving the accuracy of subsequent processing.

[0086] S1.2. Construct a prefix dictionary of common words and perform word segmentation on the analyzed text.

[0087] Specifically, word segmentation refers to using the maximum word length segmentation algorithm based on the prefix dictionary to split the input string into independent words and obtain the word segmentation result.

[0088] It should be noted that the word segmentation algorithm based on the prefix dictionary can more effectively process complex Chinese texts, reduce errors in polysemy and phrase segmentation. Through this word segmentation method, the word segmentation quality can be significantly improved, providing a more accurate basis for subsequent processing. In the existing technology, the processing method for texts is not flexible enough and is prone to errors when dealing with polysemy and phrases, affecting the overall text understanding.

[0089] S1.3. Perform part-of-speech tagging on the text after word segmentation.

[0090] Specifically, the specific steps of part-of-speech tagging are as follows:

[0091] Select the conditional random field as the model for part-of-speech tagging;

[0092] Extract the current word features, context features, and part-of-speech features from the word segmentation result;

[0093] Pass each word segmentation and its features to the conditional random field model, and calculate the part-of-speech probability of each word. The expression is:

[0094]

[0095] Where P(y i |w i ) represents the probability of the part-of-speech tag y i given the word segmentation w i . w i represents the i-th word segmentation, y i represents the part-of-speech tag associated with the i-th word segmentation w i , i represents the word segmentation index, θ represents the parameter vector of the model, f represents the set of features, T represents the transpose operation, Y represents the set of all possible part-of-speech tags, and y' represents the tag in the part-of-speech tag set Y;

[0096] For each word segmentation, select the one with the highest probability from all possible part-of-speech tags as the part-of-speech annotation of this word segmentation.

[0097] It should be noted that through the conditional random field model, context information can be combined to provide more accurate part-of-speech tagging, improving the overall accuracy of text analysis. In the existing technology, the part-of-speech tagging process is not elaborated in detail, which may lead to inaccurate syntactic analysis and affect the effectiveness of information extraction.

[0098] S1.4. After completing part-of-speech tagging, perform named entity recognition on the text.

[0099] Specifically, select a model that combines a bidirectional long short-term memory network and a conditional random field for named entity recognition. The bidirectional long short-term memory network can effectively capture context information, and the forward and backward information flows enable the model to have a more comprehensive understanding of the context. The conditional random field is used to optimize the label sequence to ensure the coherence and accuracy of the recognition results. The specific steps are as follows.

[0100] Extract word features, part-of-speech features, and context features from the results of part-of-speech tagging; in the training stage, use the labeled dataset to train the model that combines a bidirectional long short-term memory network and a conditional random field to learn how to recognize various named entities. During the training process, the model continuously adjusts the weights through the backpropagation algorithm to improve the recognition accuracy. In the inference stage, pass the input text data (including word segmentation results and part-of-speech tagging) to the trained model. For each word in the text, the model outputs a label indicating whether the word is a named entity and its category (such as person name, place name, organization name, etc.). After entity recognition is completed, summarize the recognized entities and their categories.

[0101] Furthermore, the word feature refers to the text content of the current word, and the part-of-speech feature refers to the part-of-speech label of each word, which helps the model understand the syntactic role of the word, and the lexical information within a certain range before and after the current word to capture the impact of the context on entity recognition. For example, the words before and after the current word may affect whether the word is an entity.

[0102] It should be noted that using the model that combines a bidirectional long short-term memory network and a conditional random field for named entity recognition can comprehensively understand the context, thereby improving the accuracy of entity recognition, ensuring the coherence of the recognition results, and reducing the inconsistencies in recognition. In the existing technology, the lack of context understanding leads to a decrease in the accuracy of entity recognition and affects the quality of information retrieval.

[0103] S1.5. Combine the results of word segmentation, part-of-speech tagging, and named entity recognition to perform keyword extraction and topic recognition; combine the extracted keywords, user intentions, and their related attributes into a structured query vector and store it in the knowledge base.

[0104] Specifically, the specific steps of keyword extraction and topic recognition are as follows.

[0105] The TF-IDF (Term Frequency-Inverse Document Frequency) algorithm is adopted to calculate the TF-IDF value of each word, and the expression is,

[0106] TF-IDF(t, d) = TF(t, d) × IDF(t);

[0107] Among them, TF-IDF(t, d) represents the TF-IDF value of word t in document d, t represents a specific word, d represents a specific document, TF(t, d) represents the term frequency of word t in document d, and IDF(t) represents the inverse document frequency of word t;

[0108] All words are sorted according to the TF-IDF value, and the top N words with the highest TF-IDF values are selected as keywords; by analyzing the relationships and contexts among the keywords, the main theme of the text is identified.

[0109] It should be noted that using the TF-IDF algorithm for keyword extraction can quantify the importance of each word, thereby effectively identifying the text theme and providing a structured query vector, while the keyword extraction process in the prior art is too simple and fails to fully utilize the important information in the text.

[0110] S2. Based on the structured query vector, constructing a knowledge graph to obtain a rich semantic context graph includes the following steps.

[0111] Connect to the selected knowledge base through a database connection tool; extract relevant metadata according to the user's query conditions; perform keyword matching in the knowledge base based on the extracted keywords and themes to identify multiple entities related to the user's query; retrieve the relationships between these entities; classify the obtained relationships and extract relevant attributes to judge the interactions between entities; create a node for each identified entity and create an edge for the identified relationships between entities; construct the identified entities and their relationships into a knowledge graph; for each node in the knowledge graph, initialize its feature vector; introduce the message passing mechanism of the graph neural network, and the node passes its feature vector to adjacent nodes; set the number of inference layers K, and through multi-layer graph neural network inference, gradually update the feature vector of each node by combining the feature vectors of its neighbor nodes to obtain the final feature vector of each node; analyze the final feature vector of each node to identify other themes and concepts similar to the user's query intention and construct a rich semantic context graph.

[0112] Specifically, the metadata includes the basic information of the entity (such as name, type, description, etc.) and related attributes and relationships.

[0113] Specifically, for the situation where the same keyword may correspond to multiple entities, disambiguation processing is required.

[0114] Furthermore, disambiguation refers to the situation in the knowledge base where the same keyword may correspond to multiple entities. The system needs to identify these potential ambiguities to ensure the accuracy of subsequent operations. The system will analyze the attributes and relevant information of the identified entities. By comparing these attributes, the system can determine the entity most relevant to the query intent. In addition to attribute analysis, the system will also consider context information. Based on the complete sentence entered by the user or the query background, the system can more accurately determine the most suitable entity. After attribute and context analysis, the system will select the most relevant entity as the final result.

[0115] It should be noted that based on the extracted keywords and topics, keyword matching is performed to identify multiple entities related to the user's query. Through dynamic matching, the system can quickly identify the entity most relevant to the user's intent, improving the accuracy of information retrieval, while the existing technology mainly focuses on static extraction; through relationship classification and attribute extraction, the interaction between entities is clarified, providing a deep semantic background for the construction of the knowledge graph, while the existing technology does not design in-depth analysis of entity relationships and lacks understanding of information relevance; a node is created for each identified entity, and edges are created according to the relationships between entities to form a complete knowledge graph, which can intuitively display entities and their relationships, helping users better understand the association between information, while the existing technology does not mention the construction of the knowledge graph and mainly relies on simple information extraction; through the reasoning ability of the graph neural network, deeper semantic understanding can be achieved, and other topics and concepts similar to the user's query intent can be identified, while the existing technology does not mention the application of the graph neural network and lacks in-depth analysis of complex relationships; for the situation where the same keyword may correspond to multiple entities, disambiguation is performed. Through attribute and context analysis, the most relevant entity is judged to ensure the accuracy of the final result, improve the user experience, and avoid confusion caused by ambiguity, while the existing technology does not mention the disambiguation mechanism, resulting in ambiguity in information retrieval.

[0116] S3. Introduce knowledge transfer learning technology, analyze the user's historical search data and behavior characteristics, and the steps for constructing a personalized search model are as follows.

[0117] S3.1. Extract the feature vectors of each node and the feature vectors of the edges from the semantic context graph.

[0118] Specifically, the features of each node include information such as the name, type, description, and relevance score of the node. Word embedding or graph embedding technology is used to generate these feature vectors to ensure that they can effectively capture the semantic information of the node; extract the feature vectors of the edges, which reflect the relationships between nodes. Edge features include relationship type, weight (indicating the strength of the relationship), and context information.

[0119] S3.2. Integrate the extracted node feature vectors and edge feature vectors into an input matrix.

[0120] Specifically, the structure of the input matrix is such that each row represents a node or an edge, and each column represents the extracted features.

[0121] S3.3. Collect the user's historical search data and behavioral characteristics and perform feature extraction to form a user feature matrix; concatenate the input feature matrix and the user feature matrix to form a comprehensive feature matrix.

[0122] Specifically, the user's historical search data includes query terms, clicked results, dwell time, feedback information, etc., which help to understand the user's preferences and needs.

[0123] Extract features from the user's behavior, such as:

[0124] Query frequency: The frequency of the user's searches for a specific topic.

[0125] Click-through rate: The proportion of the user's clicks on specific search results.

[0126] Dwell time: The time the user stays on a specific result page.

[0127] User feedback: The user's satisfaction rating for the search results (such as liking or disliking).

[0128] S3.4. Construct an adjacency matrix to represent the connection relationships between nodes in the graph; select a transfer learning model suitable for processing graph-structured data, and input the comprehensive feature matrix and the adjacency matrix into the transfer learning model to fine-tune the model and construct a personalized search model.

[0129] Specifically, create an adjacency matrix A to represent the connection relationships between nodes in the graph. The element A of the adjacency matrix mn can represent the relationship strength between node m and node n. If two nodes are connected, then A mn = 1, otherwise 0. Through the adjacency matrix, the model can understand the structural information between nodes.

[0130] It should be noted that creating an adjacency matrix to represent the connection relationships between nodes in the graph, and inputting the comprehensive feature matrix and the adjacency matrix into the transfer learning model for fine-tuning helps the model understand the structural information between nodes, improves the model's ability to handle complex relationships, and at the same time the fine-tuned model can better adapt to the user's personalized needs. However, the construction process of the adjacency matrix is not mentioned in the prior art, resulting in insufficient understanding of node relationships.

[0131] S4. Monitor the user's real-time interaction behavior, judge the user's cognitive load, and generate an adaptive search interface, including the following steps

[0132] Collect the interaction behavior data of users on the personalized search result page in real time; evaluate the cognitive load of users according to the interaction behavior data of users; formulate a dynamic interface adjustment strategy based on the evaluation result of users' cognitive load; adjust the layout of the interface for different cognitive load levels to generate an adaptive search interface.

[0133] Furthermore, the specific strategies include:

[0134] Simplify options: When the cognitive load is too high, the system will hide or combine some unimportant options to reduce the amount of information that users need to process.

[0135] Highlight key information: When the cognitive load is moderate, the system will help users quickly obtain key information by highlighting or magnifying important information.

[0136] Provide recommendations: When the cognitive load is low, the system can provide relevant recommended options to encourage users to conduct in-depth exploration.

[0137] Furthermore, adjust the layout of the interface for different cognitive load levels. For example:

[0138] High load: Adopt a simple card layout to highlight key information and avoid visual chaos.

[0139] Medium load: Use grouping and labels to organize information to ensure that users can quickly find the content they need.

[0140] Low load: Display more information and options to encourage users to further explore.

[0141] It should be noted that by using the collected interaction behavior data and evaluating the cognitive load level of users through an analysis algorithm, the psychological state of users at a specific moment can be understood, and corresponding adjustments can be made. However, in the existing technology, there is a lack of dynamic evaluation of users' cognitive load and the psychological changes of users during use are not considered; according to the evaluation result of cognitive load, formulate corresponding interface adjustment strategies to meet the needs of users, and adjustment strategies can be formulated in advance for different cognitive load levels to improve the response speed of the interface and user experience. However, in the existing technology, the formulation of dynamic interface adjustment strategies is not mentioned, which may lead to the rigidity of interface design; according to different cognitive load levels, adjust the layout of the interface to generate an adaptive search interface. Through reasonable interface layout adjustment, the system can effectively improve users' information processing ability, reduce cognitive load, and improve search efficiency. However, in the existing technology, the impact of cognitive load on interface design is not considered, resulting in difficulties for users in information processing.

[0142] S5. The steps of using a personalized search model to initiate multiple query requests, and after filtering and sorting, outputting a personalized search result set include the following,

[0143] S5.1. Analyze the user's query intention using a personalized search model to generate multiple query requests.

[0144] The specific operations include:

[0145] Intention recognition: The model generates relevant keywords and phrases based on the user's current query and historical behavior.

[0146] Query expansion: Expand the recognized intention to generate multiple variant queries. For example, if the user queries "artificial intelligence", relevant queries such as "AI applications", "machine learning techniques", and "deep learning" can be generated.

[0147] It should be noted that the model generates relevant keywords and phrases based on the user's current query and historical behavior, accurately identifies the user's intention, making subsequent queries more targeted, enhancing the relevance of search results. Expanding the recognized intention to generate multiple variant queries can cover a wider search range, thereby improving the comprehensiveness of information retrieval. In the existing technology, only a single query input by the user is relied on, lacking in-depth analysis and expansion of the user's intention.

[0148] S5.2. Send the generated query requests to the knowledge base and relevant databases to capture information related to the query requests; according to the user's personalized characteristics and the evaluation results of cognitive load, preliminarily filter the captured results to remove irrelevant information; sort the filtered results, evaluate the diversity of the results, and generate the final search result set.

[0149] Specifically, diversity evaluation refers to evaluating the diversity of the results in the sorted result set to ensure meeting the different needs of users. For example, the following diversity metrics can be used:

[0150] Coverage rate: Calculate the proportion of results of different topics or types.

[0151] Degree of aggregation: Evaluate the similarity between results to ensure the diversity between different results.

[0152] It should be noted that by sending the generated query requests to the knowledge base and related databases and retrieving information related to the query requests, a large number of potential results can be obtained, providing a rich basis for subsequent filtering and sorting. According to the user's personalized characteristics and the evaluation results of cognitive load, irrelevant information is removed, ensuring the relevance of the search results, avoiding interference from irrelevant information to the user, enhancing the effectiveness of the information, sorting the filtered results, evaluating the diversity of the results, generating the final search result set, ensuring that the search results are not only relevant but also meet the different needs of the user, and enhancing the richness of the information. In the prior art, the evaluation of result diversity is not mentioned, resulting in the search results being possibly too single and unable to meet the comprehensive needs of the user.

[0153] This embodiment also provides an accurate search system for Internet massive data based on AI technology, including:

[0154] A query input module that receives the query text input by the user and uses natural language processing technology to identify the input query text and generate a structured query vector;

[0155] A knowledge graph construction module that constructs a knowledge graph based on the structured query vector to obtain a rich semantic context graph;

[0156] A knowledge transfer learning module that introduces knowledge transfer learning technology, analyzes the user's historical search data and behavioral characteristics, and constructs a personalized search model;

[0157] A cognitive load evaluation module that monitors the user's real-time interaction behavior, judges the user's cognitive load, and generates an adaptive search interface;

[0158] A query request generation module that uses the personalized search model to initiate multiple query requests, and after filtering and sorting, outputs a personalized search result set.

[0159] This embodiment also provides a computer device applicable to the case of the accurate search method for Internet massive data based on AI technology, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the accurate search method for Internet massive data based on AI technology as proposed in the above embodiment.

[0160] The computer device may be a terminal, which includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or buttons, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0161] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for precise search of Internet mass data based on AI technology as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.

[0162] In summary, the present invention realizes a deep understanding of the user's query intention by receiving the query text input by the user and using natural language processing technology to generate a structured query vector; constructs a knowledge graph and generates a rich semantic context graph, enabling the system to identify entities and their multiple relationships, thereby providing more relevant search results and promoting the user to discover potential needs; introduces knowledge transfer learning technology, constructs a personalized search model by analyzing the user's historical search data, realizes precise recommendation for individual users, improves the relevance of the search, and reduces the user's cognitive load; dynamically optimizes the user experience by generating an adaptive search interface through real-time monitoring of the user's interaction behavior, making information retrieval more efficient and user-friendly; uses the personalized search model to initiate multiple query requests, filters and sorts them, and outputs a relevant and diverse set of search results to ensure that the user can obtain information efficiently.

[0163] Example 2. Referring to Table 1, this is the second embodiment of the present invention. To further verify the technical solution of the present invention, experimental simulation data of an accurate search method for a large amount of Internet data based on AI technology is given.

[0164] First, prepare a dataset containing various user query texts, and the text content covers different fields such as technology, health, and tourism. Each text will go through the following processing steps:

[0165] Simulate the query text input by the user in the search interface, use a dependency parsing tool to analyze the input text, identify the words and their parts of speech in the sentence, generate a dependency relationship graph, construct a prefix dictionary and adopt a maximum word length tokenization algorithm to tokenize the analyzed text; select a conditional random field model for part-of-speech tagging, extract the features of each token, and calculate its part-of-speech probability to determine the part-of-speech tag of each word; use a model that combines bidirectional long short-term memory network and conditional random field to perform named entity recognition on the tokens after part-of-speech tagging; adopt the TF-IDF algorithm to calculate the TF-IDF value of each word, extract keywords, and determine the main theme of the text by analyzing the relationship between the keywords.

[0166] Specifically, it is shown in Table 1 below:

[0167] Table 1 Record table of comparison data in query processing

[0168]

[0169]

[0170] Through the analysis of the above experimental data, it can be clearly seen the significant advantages of the present invention in processing user queries. The specific analysis is as follows:

[0171] For the query on how to improve immunity, using the method of the present invention only takes 0.45 seconds, while the prior art takes 1.20 seconds; when processing the recommendation of the best tourist destinations, the processing accuracy rate of the present invention reaches 92%, while the prior art is only 65%; for the query on healthy diet suggestions, 5 keywords are extracted, while the prior art only extracts 2 keywords. This shows that the present invention has a higher performance in the ability of text analysis and information extraction.

[0172] In terms of named entity recognition, the present invention has successfully recognized more entities, further enhancing the understanding of user intentions. This is crucial for constructing accurate search results.

[0173] The accuracy rate of the present invention in user intention recognition reaches 90%, while the prior art is 60%. This shows that the present invention can better meet user needs and provide a personalized search experience.

[0174] In summary, the present invention performs excellently in all aspects of processing user queries, showing its innovation in technology and great potential in practical applications. Through further optimization and application, this method is expected to provide a more accurate and efficient solution in the field of Internet data search.

[0175] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. An accurate search method for massive Internet data based on AI technology, characterized in that: including Receiving the query text input by the user and using natural language processing technology to identify the input query text, and generating a structured query vector includes the following steps The user inputs the query text in the search interface Using a dependency syntax analysis tool to perform basic syntax analysis on the input text Constructing a prefix dictionary of common words and performing word segmentation on the analyzed text Performing part-of-speech tagging on the text after word segmentation After completing part-of-speech tagging, performing named entity recognition on the text Combining the results of word segmentation, part-of-speech tagging, and named entity recognition to perform keyword extraction and topic recognition Combining the extracted keywords, user intentions, and their related attributes into a structured query vector and storing it in the knowledge base Based on the structured query vector, constructing a knowledge graph to obtain a rich semantic context graph includes the following steps Connecting to the selected knowledge base through a database connection tool Extracting relevant metadata according to the user's query conditions Based on the extracted keywords and topics, performing keyword matching in the knowledge base to identify multiple entities related to the user's query According to the identified multiple related entities, retrieving the relationships between these entities Classifying the obtained relationships and extracting relevant attributes to judge the interaction between entities Creating nodes for each identified entity and identifying the relationships created between entities Constructing the identified entities and their relationships into a knowledge graph For each node in the knowledge graph, initializing its feature vector Introducing the message passing mechanism of the graph neural network, and the node passes its feature vector to adjacent nodes Setting the number of inference layers K, and through multi-layer graph neural network inference, gradually updating the feature vector of each node by combining the feature vectors of its neighbor nodes to obtain the final feature vector of each node Analyzing the final feature vector of each node, identifying other topics and concepts similar to the user's query intention, and constructing a rich semantic context graph Introducing knowledge transfer learning technology, analyzing the user's historical search data and behavioral characteristics, and constructing a personalized search model Monitoring the user's real-time interaction behavior, judging the user's cognitive load, and generating an adaptive search interface Using the personalized search model to initiate multiple query requests, and after filtering and sorting, outputting a personalized search result set 2. The precise search method for Internet mass data based on AI technology according to claim 1, characterized in that: The basic syntax analysis refers to using a dependency syntax analysis tool to parse the input text, identifying each word in the sentence and its corresponding part of speech, automatically analyzing the relationships between words, and generating a dependency relationship graph; the word segmentation process refers to using the maximum word length segmentation algorithm based on the prefix dictionary to split the input string into independent words to obtain the word segmentation result The specific steps of the part-of-speech tagging are as follows Selecting a conditional random field as the model for part-of-speech tagging Extracting the current word features, context features, and part-of-speech features from the word segmentation result Passing each word segmentation and its features to the conditional random field model to calculate the part-of-speech probability of each word, and the expression is Among them, P(y i | w i ) represents the probability of the part-of-speech tag y i given the word segmentation w i , w i represents the i-th word segmentation, y i represents the part-of-speech tag associated with the i-th word segmentation w i , i represents the word segmentation index, θ represents the parameter vector of the model, f represents the set of features, T represents the transpose operation, Y represents the set of all possible part-of-speech tags, and y' represents a tag in the part-of-speech tag set Y; For each word segmentation, selecting the one with the highest probability from all possible part-of-speech tags as the part-of-speech tagging of the word segmentation The process of the named entity recognition is as follows Select a model that combines a bidirectional long short-term memory network and a conditional random field for named entity recognition; Input the word segmentation with part-of-speech tagging into the model for training; Pass the word segmentation result and part-of-speech tagging to the trained model, perform named entity recognition on each word, and output the corresponding entity category; The specific steps of the above keyword extraction and topic recognition are as follows, Use the TF-IDF algorithm to calculate the TF-IDF value of each word. The expression is, TF-IDF(t, d) = TF(t, d) × IDF(t); Among them, TF-IDF(t, d) represents the TF-IDF value of word t in document d, t represents a specific word, d represents a specific document, TF(t, d) represents the word frequency of word t in document d, and IDF(t) represents the inverse document frequency of word t; Sort all words according to the TF-IDF value, and select the top N words with the highest TF-IDF value as keywords; Identify the main theme of the text by analyzing the relationships and contexts among the keywords.

3. The precise search method for Internet mass data based on AI technology according to claim 2, wherein: Introduce knowledge transfer learning technology, analyze the user's historical search data and behavior characteristics, and build a personalized search model, including the following steps, Extract the feature vectors of each node and the feature vectors of the edges from the semantic context graph; Integrate the extracted node feature vectors and edge feature vectors into an input matrix; Collect the user's historical search data and behavior characteristics and perform feature extraction to form a user feature matrix; Concatenate the input feature matrix and the user feature matrix together to form a comprehensive feature matrix; Construct an adjacency matrix to represent the connection relationships between the nodes in the graph; Select a transfer learning model suitable for processing graph-structured data, input the comprehensive feature matrix and the adjacency matrix into the transfer learning model to fine-tune the model, and build a personalized search model.

4. The precise search method for Internet mass data based on AI technology according to claim 3, wherein: Monitor the user's real-time interaction behavior, judge the user's cognitive load, and generate an adaptive search interface, including the following steps, Collect the interaction behavior data of the user on the personalized search result page in real time; Evaluate the user's cognitive load according to the user's interaction behavior data; Formulate a dynamic interface adjustment strategy based on the evaluation result of the user's cognitive load; Adjust the layout of the interface for different cognitive load levels to generate an adaptive search interface.

5. The precise search method for Internet mass data based on AI technology according to claim 4, characterized in that: Use the personalized search model to initiate multiple query requests, and after filtering and sorting, output a personalized search result set, including the following steps, Use the personalized search model to analyze the user's query intent and generate multiple query requests; Send the generated query requests to the knowledge base and related databases to retrieve information related to the query requests; Perform preliminary filtering on the retrieved results according to the user's personalized characteristics and the evaluation result of the cognitive load; Sort the filtered results, evaluate the diversity of the results, and generate the final search result set.

6. An Internet massive data precise search system based on AI technology, based on the Internet massive data precise search method based on AI technology according to any one of claims 1 to 5, characterized in that: Including, A query input module that receives the query text input by the user and uses natural language processing technology to recognize the input query text and generate a structured query vector; A knowledge graph construction module that constructs a knowledge graph based on the structured query vector to obtain a rich semantic context graph; A knowledge transfer learning module, which introduces knowledge transfer learning technology, analyzes the user's historical search data and behavioral characteristics, and constructs a personalized search model; A cognitive load assessment module, which monitors the user's real-time interaction behavior, judges the user's cognitive load, and generates an adaptive search interface; A query request generation module, which uses the personalized search model to initiate multiple query requests, and after filtering and sorting, outputs a personalized search result set.

7. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the precise Internet mass data search method based on AI technology according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the precise Internet mass data search method based on AI technology according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Self-adaptive human-computer interface configuration method

    CN109558005A

  • Satellite image recommendation method based on user demand understanding

    CN115374303A