Data Intelligent Processing Method Based on Artificial Intelligence Search Engine Technology
By combining the target user's historical search records and real-time web page information in the artificial intelligence search engine, semantic joint and parameter updates are solved, and the problem of inaccurate output results of large language models is achieved, achieving more efficient and accurate search result output.
Patent Information
- Application Number
- CN202510703000.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-29
AI Technical Summary
In the existing artificial intelligence search engine technology, the output results of large language models are poor in accuracy and real-time, especially when facing dynamic data scenarios, traditional parameter adjustment costs high and it is difficult to ensure accuracy.
By extracting keywords and web page information in the target user's historical search record, semantic jointing is performed, the frequency of joint semantics and web page information differences are calculated, and the parameters of the target model are updated to improve the accuracy and real-timeness of the output results.
It improves the search efficiency of search engines and the directional accuracy of output results, ensures real-time update and accuracy of model parameters, and adapts to dynamic data scenarios.
Smart Images

Figure CN120234403B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a data intelligent processing method based on artificial intelligence search engine technology. Background Art
[0002] An artificial intelligence search engine is a new generation of search system that combines information retrieval and large language models, such as DeepSeek, Bing Chat, Wolfram Alpha, etc. It uses technologies such as natural language processing (semantic understanding, such as distinguishing between "Apple Inc." and "fruit apple"), large language models (such as GPT-4, Gemini, Claude, etc., for generating answers and summarizing information), and knowledge graphs (structuring and associating data such as people, events, locations, etc. to provide direct answers) to improve the intelligence, accuracy, and interaction experience of searches: users input search content into a pre-trained large language model, the natural language processing technology generates search keywords or query statements based on the search content, the search engine retrieves information according to the search keywords or query statements, and returns the retrieved results to the large language model. After integrating and analyzing the received results, the large language model finally outputs the results.
[0003] For large language models, some parameters in traditional large language models need to be predefined manually. However, with the change of data scenarios, the manually predefined parameters cannot guarantee the accuracy of the data processing results of the large language model. For example, in a financial evaluation model, as factors such as the market environment and policies and regulations change continuously, traditional financial evaluation models need to be adjusted manually constantly, resulting in a high maintenance cost for the financial evaluation model and it is difficult to guarantee the accuracy of the data processing of the financial evaluation model. At the same time, in terms of information retrieval, traditional search engines mainly rely on keyword matching and cannot rely on historical search records, resulting in poor semantic understanding ability and unable to output the results required by users, that is, the accuracy of the output results is poor, which in turn leads to a poor user experience. For example, when a user inputs "apple", the search engine cannot accurately determine whether the user wants to know information about the fruit "apple" or the product dynamics of the technology company "Apple", resulting in the search engine returning a large amount of content irrelevant to the user's actual needs to the large language model, resulting in poor directional accuracy of the output content of the large language model, and the user needs to input a qualifier again to obtain valuable information.
[0004] Therefore, how to improve the accuracy of the output results of large language models in artificial intelligence search engine technology has become an urgent problem to be solved. Summary of the Invention
[0005] In view of this, an embodiment of the present invention provides a data intelligent processing method based on artificial intelligence search engine technology to solve the problem of how to improve the accuracy of the output results of large language models in artificial intelligence search engine technology.
[0006] An embodiment of the present invention provides a data intelligent processing method based on artificial intelligence search engine technology, and the method includes the following steps:
[0007] According to the search content input by the target user in the target model at the current moment, at least one keyword is extracted. For any keyword, semantic association is performed on the any keyword according to the historical search records of the target user in the target model to obtain at least one associated semantics of the any keyword. The historical search records include historical keywords corresponding to the historical search content of the target user and historical web page information corresponding to each historical keyword;
[0008] According to the occurrence frequency of the historical keywords in each associated semantics in the historical search records, the association coefficient of each associated semantics is obtained. According to each historical web page information containing the associated semantics in the historical search records and the association coefficient of each associated semantics, the target associated semantics of the any keyword are obtained;
[0009] According to the target associated semantics of each keyword, the real-time web page information of each target associated semantics is obtained. According to the difference between the historical web page information and the real-time web page information of each target associated semantics, the parameters in the target model are updated to obtain an updated target model. According to the updated target model, the search results corresponding to the search content are output.
[0010] Preferably, the obtaining the association coefficient of each associated semantics according to the occurrence frequency of the historical keywords in each associated semantics in the historical search records includes:
[0011] For any associated semantics, in the historical search records, at least one target search record of the historical keywords in the any associated semantics and the time stamp of each target search record are obtained. The time interval between the time stamp of each target search record and the current moment is calculated, and the reciprocal of each time interval is linearly normalized to obtain the time feature value of each target search record;
[0012] Calculate the proportion of the number of all target search records in the number of all historical search records to obtain the search frequency of the historical keywords in the any associated semantics;
[0013] Perform weighted summation on the search frequency and the average value of the time feature values of all target search records to obtain the association coefficient of the any associated semantics.
[0014] Preferably, obtaining the target combined semantics of any keyword according to each piece of historical web page information containing the combined semantics in the historical search record and the combination coefficient of each combined semantics includes:
[0015] For any combined semantics, in the historical search record, obtain the number of pieces of historical web page information containing the any combined semantics to obtain the number of web page information of the any combined semantics, obtain the number of web page information of all combined semantics of the any keyword, calculate the proportion of the number of web page information of the any combined semantics in the number of web page information of all combined semantics of the any keyword to obtain the web page frequency of the any combined semantics;
[0016] Perform linear normalization on the web page frequency of the any combined semantics to obtain a normalized web page frequency value, calculate the product between the normalized web page frequency value of the any combined semantics and the combination coefficient to obtain the expected possible probability of the any combined semantics;
[0017] Obtain the expected possible probabilities of all combined semantics of the any keyword, and obtain the target combined semantics of the any keyword according to the expected possible probabilities of all combined semantics of the any keyword.
[0018] Preferably, obtaining the target combined semantics of the any keyword according to the expected possible probabilities of all combined semantics of the any keyword includes:
[0019] Among the expected possible probabilities of all combined semantics of the any keyword, select the combined semantics corresponding to the maximum value as the target combined semantics of the any keyword.
[0020] Preferably, updating the parameters in the target model according to the difference between the historical web page information and the real-time web page information of each target combined semantics to obtain an updated target model includes:
[0021] For any real-time web page information, in the historical search record, obtain the target historical web page information with the same website address as the any real-time web page information at different timestamps, and form the any real-time web page information and the target historical web page information into web page information to be analyzed;
[0022] According to the timestamp corresponding to each piece of web page information to be analyzed, perform text comparison on two pieces of web page information to be analyzed corresponding to every two adjacent timestamps to obtain a comparison file of two pieces of web page information to be analyzed corresponding to every two adjacent timestamps, and obtain the information difference degree of each comparison file according to the distribution of the difference data in each comparison file;
[0023] According to the information difference degrees of all the comparison documents, determine whether there is data change in any of the real-time web page information. If there is data change in any of the real-time web page information, record the real-time web page information as the updated data source web page information of the target model;
[0024] Obtain all the updated data source web page information, and update the parameters in the target model according to all the updated data source web page information to obtain the updated target model.
[0025] Preferably, the obtaining the information difference degree of each comparison document according to the distribution of the difference data in each comparison document includes:
[0026] For any comparison document, if there is no difference data in the comparison document, set the information difference degree of the comparison document to 0;
[0027] If there is difference data in any comparison document, subtract the reciprocal of the number of all difference data in the comparison document from 1 to obtain the first difference degree of the comparison document;
[0028] In any comparison document, obtain the number of characters between every two adjacent difference data, and subtract the reciprocal of the average value of all the number of characters from 1 to obtain the second difference degree of the comparison document;
[0029] Calculate the average value between the first difference degree and the second difference degree to obtain the information difference degree of the comparison document.
[0030] Preferably, the determining whether there is data change in any of the real-time web page information according to the information difference degrees of all the comparison documents includes:
[0031] Among all the information difference degrees, obtain the occurrence times of each information difference degree, and use the information difference degree corresponding to the maximum occurrence times as the information difference degree threshold;
[0032] If the information difference degree of the comparison document corresponding to any real-time web page information is greater than the information difference degree threshold, determine that there is data change in the real-time web page information;
[0033] If the information difference degree of the comparison document corresponding to any real-time web page information is less than or equal to the information difference degree threshold, determine that there is no data change in the real-time web page information.
[0034] Preferably, after obtaining all the updated data source web page information, it further includes:
[0035] If there is no updated data source web page information, output the search result corresponding to the search content according to the target model.
[0036] The beneficial effects of the embodiments of the present invention compared with the prior art are as follows:
[0037] According to the search content input by the target user in the target model at the current moment, the present invention extracts at least one keyword. For any keyword, according to the historical search records of the target user in the target model, semantic association is performed on the any keyword to obtain at least one associated semantics of the any keyword. The historical search records include historical keywords corresponding to the historical search content of the target user and historical web page information corresponding to each historical keyword; according to the occurrence frequency of the historical keywords in each associated semantics in the historical search records, the association coefficient of each associated semantics is obtained, and according to each historical web page information containing the associated semantics in the historical search records and the association coefficient of each associated semantics, the target associated semantics of the any keyword are obtained; according to the target associated semantics of each keyword, the real-time web page information of each target associated semantics is obtained, and according to the difference between the historical web page information and the real-time web page information of each target associated semantics, the parameters in the target model are updated to obtain an updated target model, and according to the updated target model, the search results corresponding to the search content are output. Among them, by combining the historical search records of the target user, semantic association is performed on the content searched by the target user at the current moment to obtain target associated semantics that better meet the user's needs, improve the search efficiency of the search engine, and further improve the directional accuracy of the content output by the target model; at the same time, according to the real-time web page information of each target associated semantics, the parameters in the target model are updated in real time to improve the accuracy and real-time performance of the output results of the target model. Brief Description of the Drawings
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0039] Figure 1 It is a method flow chart of a data intelligent processing method based on artificial intelligence search engine technology provided by Embodiment 1 of the present invention. Detailed Embodiments
[0040] The following details the embodiments of the present disclosure, and the examples of the embodiments are shown in the drawings. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present disclosure and should not be construed as a limitation of the present disclosure.
[0041] It should be noted that the terms "first", "second", etc. in the specification of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure.
[0042] In order to illustrate the technical solution of the present invention, the following will be described through specific embodiments.
[0043] See Figure 1 , which is a method flowchart of a data intelligent processing method based on artificial intelligence search engine technology provided in the first embodiment of the present invention. As Figure 1 shown, the method may include:
[0044] Step S101, according to the search content input by the target user in the target model at the current moment, extract at least one keyword. For any keyword, according to the historical search records of the target user in the target model, perform semantic association on the any keyword to obtain at least one associated semantics of the any keyword. The historical search records include historical keywords corresponding to the historical search content of the target user, and historical web page information corresponding to each historical keyword.
[0045] The search content input by the target user in the target model at the current moment is segmented by the Jieba library in Python to obtain at least one keyword. Among them, the target model is a model constructed according to the content in a fixed web page, such as a financial risk assessment model, which is constructed by financial data such as tax rates and exchange rates in a fixed official web page to perform financial risk assessment according to the input content of the target user. The Jieba library is a prior art and will not be elaborated here. After obtaining the keywords, each keyword is retrieved through a search engine. Considering that there may be a situation where the semantics of the keywords are not unique. For example, when the user inputs "apple", the search engine cannot accurately determine whether the user wants to know information related to the fruit "apple", resulting in the search engine possibly returning a large amount of data irrelevant to the needs of the target user to the target model, making the output result of the target model inaccurate.
[0046] Therefore, in the embodiments of the present invention, before retrieving each keyword through a search engine, according to the historical search records of the target user in the target model (including the historical keywords corresponding to the historical search content of the target user and the historical web page information corresponding to each historical keyword), the ColBERT model is used to perform semantic combination on each keyword to obtain at least one combined semantics for each keyword. Taking the k-th keyword as an example, if the k-th keyword is "apple", and there are historical keywords such as "model" and "memory" in the historical search records of the target user in the target model, then the combined semantics of "apple" are "apple model", "apple memory", etc., so as to analyze each combined semantics, initially determine the semantics required by the user, improve the search efficiency of the search engine, and further improve the directional accuracy of the content output by the target model. Among them, the ColBERT model is a prior art and will not be elaborated here.
[0047] Step S102: Obtain the combined coefficient of each combined semantics according to the occurrence frequency of the historical keywords in each combined semantics in the historical search records, and obtain the target combined semantics of any keyword according to each historical web page information containing the combined semantics in the historical search records and the combined coefficient of each combined semantics.
[0048] In order to improve the accuracy of judging the semantics required by the user, in the embodiments of the present invention, according to the occurrence frequency of the historical keywords in each combined semantics of the k-th keyword in the historical search records, the combined coefficient of each combined semantics of the k-th keyword is obtained, which is used to represent the possibility that the i-th combined semantics is the semantics required by the user. Taking the i-th combined semantics of the k-th keyword as an example, if the occurrence frequency of the historical keyword in the i-th combined semantics in the historical search records is relatively high, then the possibility that the i-th combined semantics is the semantics required by the user is relatively large. The specific method for obtaining the combined coefficient of the i-th combined semantics is as follows:
[0049] In the historical search records, obtain at least one target search record of the historical keyword in the i-th combined semantics and the time stamp of each target search record, calculate the time interval between the time stamp of each target search record and the current moment, and perform linear normalization on the reciprocal of each time interval to obtain the time feature value of each target search record. Among them, linear normalization is a prior art and will not be elaborated here;
[0050] Calculate the proportion of the number of all target search records in the number of all historical search records to obtain the search frequency of the historical keyword in the i-th combined semantics;
[0051] Perform weighted summation on the search frequency and the average value of the time feature values of all target search records to obtain the combined coefficient of the i-th combined semantics.
[0052] In one embodiment, the calculation formula for the combination coefficient of the i-th combined semantics is as follows:
[0053]
[0054] Wherein, represents the combination coefficient of the i-th combined semantics, represents the first weight, n represents the number of all target search records, represents the time interval between the j-th target search record and the current moment, represents the second weight, represents the search frequency of the historical keywords in the i-th combined semantics, and norm() represents the linear normalization function.
[0055] It should be noted that, the smaller, the larger, it indicates that the target user searches the historical keywords in the i-th combined semantics multiple times in a short period of time, and thus the larger, the greater the possibility that the i-th combined semantics is the semantics required by the user. Since there may be a situation where the target user searches the historical keywords in the i-th combined semantics more frequently in the early stage, the search frequency of the historical keywords in the i-th combined semantics is used as the main reference, and is set, There is no limitation here, and the implementer can set it according to the specific scenario.
[0056] Since semantic combination is an irregular combination, there may be incorrect combined semantics. If only analyzing based on the historical keywords in the i-th combined semantics, it is impossible to determine whether the i-th combined semantics itself is correct. Therefore, it is necessary to determine whether the i-th combined semantics is correct according to the occurrence situation of the i-th combined semantics in the actual data. In the embodiments of the present invention, the i-th combined semantics is searched in all historical web page information in the historical search records of the target user through the Whoosh library, and the number of historical web page information containing the i-th combined semantics is obtained, which is used to characterize the occurrence situation of the i-th combined semantics in the actual data. Then, in combination with the combination coefficient of the i-th combined semantics, the expected possible probability of the i-th combined semantics is obtained, which is used to characterize the possibility that the i-th combined semantics is the semantics required by the user. Among them, the Whoosh library is a prior art and will not be elaborated here. The specific method for obtaining the expected possible probability of the i-th combined semantics is as follows:
[0057] In the historical search records, obtain the number of historical web page information containing the i-th combined semantics, to get the number of web page information of the i-th combined semantics, obtain the number of web page information of all combined semantics of the k-th keyword, calculate the proportion of the number of web page information of the i-th combined semantics in the number of web page information of all combined semantics of the k-th keyword, to get the web page frequency of the i-th combined semantics;
[0058] Perform linear normalization on the web page frequency of the i-th combined semantics to obtain a normalized web page frequency value, calculate the product between the normalized web page frequency value of the i-th combined semantics and the combined coefficient, to get the expected possible probability of any combined semantics.
[0059] In one embodiment, the calculation formula for the expected possible probability of the i-th combined semantics is:
[0060]
[0061] Wherein, represents the expected possible probability of the i-th combined semantics, represents the combined coefficient of the i-th combined semantics, represents the web page frequency of the i-th combined semantics, and norm() represents the linear normalization function.
[0062] It should be noted that, the larger it is, it indicates that the number of occurrences of the i-th combined semantics in the actual data is more, and thus the larger it is, the greater the possibility that the i-th combined semantics is the semantics required by the user; if , it indicates that the i-th combined semantics is an incorrect combined semantics, that is , the i-th combined semantics must not be the semantics required by the user; the larger it is, it indicates that the target user searches for the historical keywords in the i-th combined semantics multiple times in a short period of time, and thus the larger it is, the greater the possibility that the i-th combined semantics is the semantics required by the user.
[0063] Similarly, obtain the expected possible probabilities of all combined semantics of the k-th keyword input by the target user in the target model, and among the expected possible probabilities of all combined semantics, select the combined semantics corresponding to the maximum value as the target combined semantics of the k-th keyword, for representing the semantics required by the user.
[0064] Step S103, according to the target combined semantics of each keyword, obtain the real-time web page information of each target combined semantics, update the parameters in the target model according to the difference between the historical web page information and the real-time web page information of each target combined semantics, to get an updated target model, and output the search result corresponding to the search content according to the updated target model.
[0065] According to the obtaining method of the target combined semantics of the k-th keyword in step S102, obtain the target combined semantics of each keyword input by the target user in the target model. Retrieve the real-time web page information of each target combined semantics through a search engine. The target model analyzes and integrates the information content in the real-time web page information of each target combined semantics, and outputs the result of the analysis and integration to the target user.
[0066] However, considering that some models themselves do not have real-time performance. For example, in a financial risk assessment model, some parameters in the financial risk assessment model are obtained based on data in some official fixed web pages and a well-known formula. With the adjustment of influencing factors such as economic prices and exchange rates, the data in these fixed web pages will change, and the parameters in the original financial risk assessment model may not be applicable to the current financial risk assessment, resulting in inaccurate output results of the original financial risk assessment model. Since the target model saves the web page information to the historical search record of the target user in the target model each time it uses a certain web page information, in the embodiment of the present invention, after obtaining the real-time web page information of each target combined semantics, first compare these real-time web page information with the historical web page information in the historical search record to determine whether the data in these real-time web page information has changed. If it has changed, the parameters of the target model need to be updated to improve the accuracy of the output result of the target model.
[0067] In the embodiment of the present invention, taking the h-th real-time web page information as an example, in the historical search record of the target user in the target model, obtain the target historical web page information with the same website address as the h-th real-time web page information at different timestamps. Combine the h-th real-time web page information with the target historical web page information to form the web page information to be analyzed. According to the timestamp corresponding to each web page information to be analyzed, use WinMerge to perform text comparison on two adjacent web page information to be analyzed corresponding to two timestamps. There is no limitation here, and the implementer can set the text comparison tool according to the specific scenario to obtain the comparison file of two adjacent web page information to be analyzed corresponding to two timestamps. Since the parameters of the target model have a direct relationship with digital data, in the comparison file, the highlighted data is the digital data with differences between the two web page information to be analyzed, and record the highlighted data as the difference data. Among them, using WinMerge for text comparison is a prior art and will not be elaborated here. Further, according to the distribution of the difference data in each comparison file, obtain the information difference degree of each comparison file. Specifically:
[0068] For any comparison file, if there is no difference data in the any comparison file, set the information difference degree of the any comparison file to 0;
[0069] If there are difference data in any of the comparison documents, subtract the reciprocal of the number of all difference data in any of the comparison documents from 1 to obtain the first difference degree of any of the comparison documents;
[0070] In any of the comparison documents, obtain the number of characters between every two adjacent difference data, and subtract the reciprocal of the average value of all the number of characters from 1 to obtain the second difference degree of any of the comparison documents;
[0071] Calculate the average between the first difference degree and the second difference degree to obtain the information difference degree of any of the comparison documents.
[0072] In an embodiment, taking the y-th comparison document as an example, the calculation formula for the information difference degree of the y-th comparison document is:
[0073]
[0074] where, represents the information difference degree of the y-th comparison document, represents the number of all difference data in the y-th comparison document, represents the number of characters between the s-th pair of adjacent difference data in the y-th comparison document.
[0075] It should be noted that since there may be data such as time that changes daily in web page information, these data are irrelevant to the parameters of the target model, and these data may be concentrated together with a small positional gap between them. Therefore, the smaller, the fewer the number of characters between two adjacent difference data, that is, the smaller the distance between every two adjacent difference data. Furthermore, the smaller, the less likely it is that two adjacent difference data are parameters related to the target model; the smaller, the fewer the number of difference data in the y-th comparison document, and the smaller the difference between the two web page information to be analyzed corresponding to the y-th comparison document. Furthermore, the smaller, the less likely it is that two adjacent difference data are parameters related to the target model.
[0076] Similarly, obtain the information difference degrees of all comparison documents. Among all the information difference degrees, obtain the number of occurrences of each information difference degree, and take the information difference degree corresponding to the maximum number of occurrences as the information difference degree threshold. If the information difference degree of the comparison document corresponding to the h-th real-time web page information is greater than the information difference degree threshold, it is determined that the h-th real-time web page information has changed data; if the information difference degree of the comparison document corresponding to the h-th real-time web page information is less than or equal to the information difference degree threshold, it is determined that the h-th real-time web page information has not changed data.
[0077] If the data of the h-th real-time web page information changes, the h-th real-time web page information is recorded as the web page information of the updated data source of the target model. Similarly, it is judged whether the data of all real-time web page information changes to obtain all web pages of the updated data source, and the parameters in the target model are updated according to the information of all web pages of the updated data source to obtain the updated target model. For example, an industry evaluation model is constructed. This model is based on the number of companies, company levels, the number of employees, the talent gap, the average salary, etc. of each industry, and finally obtains a comprehensive evaluation of a certain industry to provide reference for job seekers. Suppose the key built-in parameters of the industry evaluation model are: the number of companies; company levels; the number of employees; the talent gap; the average salary, etc. The target user inputs "communication" in this model, and after semantic combination of "communication", the target combined semantics is obtained as "communication software". The average salary of "communication software" in the current industry evaluation model is 10,000, but according to the information of the updated data source web page, the current average salary of "communication software" is 8,000. Then the average salary of "communication software" in the industry evaluation model is updated to 8,000 to obtain the updated industry evaluation model.
[0078] According to the updated target model, analyze and integrate the information content in the real-time web page information of each target combined semantics, and output the analyzed result to the user.
[0079] Specifically, if there is no information of the updated data source web page, it means that the parameters in the target model do not need to be updated. Then the target model analyzes and integrates the real-time web page information of each target combined semantics returned by the search engine, and outputs the analyzed and integrated result to the target user.
[0080] It should be noted that the focus of the present invention is: how to perform semantic combination according to the search content of the target user to improve the search efficiency of the search engine, thereby improving the directional accuracy of the content output by the target model. At the same time, according to the real-time web page information retrieved by the search engine, it is judged whether the parameters in the target model need to be updated, and then the parameters in the target model are updated to improve the accuracy of the output result of the target model. Among them, the update of the parameters in the target model is the prior art and will not be elaborated here.
[0081] In summary, the present invention extracts at least one keyword according to the search content input by the target user in the target model at the current moment. For any keyword, semantic combination is performed on the any keyword according to the historical search record of the target user in the target model, and at least one combined semantics of the any keyword is obtained. The historical search record includes historical keywords corresponding to the historical search content of the target user and historical web page information corresponding to each historical keyword; according to the occurrence frequency of the historical keywords in each combined semantics in the historical search record, the combined coefficient of each combined semantics is obtained, and according to each historical web page information containing the combined semantics in the historical search record and the combined coefficient of each combined semantics, the target combined semantics of the any keyword is obtained; according to the target combined semantics of each keyword, the real-time web page information of each target combined semantics is obtained, and according to the difference between the historical web page information and the real-time web page information of each target combined semantics, the parameters in the target model are updated to obtain an updated target model, and according to the updated target model, the search result corresponding to the search content is output. Among them, by combining the historical search record of the target user, semantic combination is performed on the content searched by the target user at the current moment, and more user-demand-compliant target combined semantics are obtained, improving the search efficiency of the search engine and further improving the directional accuracy of the content output by the target model; at the same time, according to the real-time web page information of each target combined semantics, the parameters in the target model are updated in real time, improving the accuracy and real-time nature of the output result of the target model.
[0082] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A data intelligent processing method based on artificial intelligence search engine technology, characterized in that The data intelligent processing method based on the artificial intelligence search engine technology includes: According to the search content input by the target user in the target model at the current moment, at least one keyword is extracted. For any keyword, according to the historical search records of the target user in the target model, semantic combination is performed on the any keyword to obtain at least one combined semantics of the any keyword. The historical search records include historical keywords corresponding to the historical search content of the target user and historical web page information corresponding to each historical keyword; According to the occurrence frequency of the historical keywords in each combined semantics in the historical search records, the combined coefficient of each combined semantics is obtained. According to each historical web page information containing the combined semantics in the historical search records and the combined coefficient of each combined semantics, the target combined semantics of the any keyword is obtained; According to the target combined semantics of each keyword, the real-time web page information of each target combined semantics is obtained. According to the difference between the historical web page information and the real-time web page information of each target combined semantics, the parameters in the target model are updated to obtain an updated target model. According to the updated target model, the search result corresponding to the search content is output.
2. The data intelligent processing method based on artificial intelligence search engine technology according to claim 1, characterized in that The obtaining of the combined coefficient of each combined semantics according to the occurrence frequency of the historical keywords in each combined semantics in the historical search records includes: For any combined semantics, in the historical search records, at least one target search record of the historical keywords in the any combined semantics and the time stamp of each target search record are obtained. The time interval between the time stamp of each target search record and the current moment is calculated, and the reciprocal of each time interval is linearly normalized to obtain the time feature value of each target search record; Calculate the proportion of the number of all target search records in the number of all historical search records to obtain the search frequency of the historical keywords in the any combined semantics; Perform weighted summation on the search frequency and the average value of the time feature values of all target search records to obtain the combined coefficient of the any combined semantics.
3. The data intelligent processing method based on artificial intelligence search engine technology according to claim 1, wherein The obtaining of the target combined semantics of the any keyword according to each historical web page information containing the combined semantics in the historical search records and the combined coefficient of each combined semantics includes: For any combined semantics, in the historical search records, the number of all historical web page information containing the any combined semantics is obtained to obtain the number of web page information of the any combined semantics. The number of web page information of all combined semantics of the any keyword is obtained. Calculate the proportion of the number of web page information of the any combined semantics in the number of web page information of all combined semantics of the any keyword to obtain the web page frequency of the any combined semantics; Linearly normalize the web page frequency of the any combined semantics to obtain a normalized web page frequency value. Calculate the product of the normalized web page frequency value of the any combined semantics and the combined coefficient to obtain the expected possible probability of the any combined semantics; Obtain the expected possible probabilities of all combined semantics of any one of the keywords, and obtain the target combined semantics of any one of the keywords according to the expected possible probabilities of all combined semantics of any one of the keywords.
4. The data intelligent processing method based on artificial intelligence search engine technology according to claim 3, wherein The obtaining the target combined semantics of any one of the keywords according to the expected possible probabilities of all combined semantics of any one of the keywords includes: Among the expected possible probabilities of all combined semantics of any one of the keywords, select the combined semantics corresponding to the maximum value as the target combined semantics of any one of the keywords.
5. The data intelligent processing method based on artificial intelligence search engine technology according to claim 1, characterized in that The updating the parameters in the target model according to the differences between the historical web page information and the real-time web page information of each target combined semantics to obtain the updated target model includes: For any real-time web page information, in the historical search records, obtain the target historical web page information with the same website address as the any real-time web page information at different timestamps, and form the web page information to be analyzed by combining the any real-time web page information and the target historical web page information; According to the timestamps corresponding to each piece of web page information to be analyzed, perform text comparison on two pieces of web page information to be analyzed corresponding to every two adjacent timestamps to obtain the comparison files of two pieces of web page information to be analyzed corresponding to every two adjacent timestamps, and obtain the information difference degree of each comparison file according to the distribution of the difference data in each comparison file; According to the information difference degrees of all comparison files, determine whether the any real-time web page information has changed. If the any real-time web page information has changed, record the any real-time web page information as the updated data source web page information of the target model; Obtain all updated data source web page information, and update the parameters in the target model according to all updated data source web page information to obtain the updated target model.
6. The data intelligent processing method based on artificial intelligence search engine technology according to claim 5, wherein The obtaining the information difference degree of each comparison file according to the distribution of the difference data in each comparison file includes: For any comparison file, if there is no difference data in the any comparison file, set the information difference degree of the any comparison file to 0; If there is difference data in the any comparison file, subtract the reciprocal of the number of all difference data in the any comparison file from the constant 1 to obtain the first difference degree of the any comparison file; In the any comparison file, obtain the number of characters between every two adjacent difference data, and subtract the reciprocal of the average value of all character numbers from the constant 1 to obtain the second difference degree of the any comparison file; Calculate the average of the first difference degree and the second difference degree to obtain the information difference degree of the any comparison file.
7. The data intelligent processing method based on artificial intelligence search engine technology according to claim 5, characterized in that, The determining whether the any real-time web page information has changed according to the information difference degrees of all comparison files includes: Among all information difference degrees, obtain the occurrence times of each information difference degree, and use the information difference degree corresponding to the maximum occurrence times as the information difference degree threshold; If the information difference degree of the comparison file corresponding to the any real-time web page information is greater than the information difference degree threshold, determine that the any real-time web page information has changed; If the information difference degree of the comparison file corresponding to any of the real-time web page information is less than or equal to the information difference degree threshold, it is determined that there is no data change in any of the real-time web page information.
8. The data intelligent processing method based on artificial intelligence search engine technology according to claim 5, characterized in that, After obtaining all the updated data source web page information, it further includes: If there is no updated data source web page information, the search result corresponding to the search content is output according to the target model.
Citation Information
Patent Citations
Website navigation service system based on artificial intelligence
CN119249011A
Keyword-based search method and device, medium and equipment
CN119557421A