Intelligent data processing method based on artificial intelligence search engine technology

By performing semantic jointing and real-time web page information updates on the historical search records of the target user, the problem of inaccurate output of large language models is solved, the efficiency of search engines and the accuracy of results are improved, and the data changes are adapted.

CN120234403AActive Publication Date: 2025-07-01SHANDONG SHENGDE INTELLIGENT TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510703000.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-01
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

In the prior art, the output results of the large language model of artificial intelligence search engines are inaccurate, resulting in poor user experience, especially in terms of financial evaluation models and information retrieval, it is difficult to ensure the accuracy and real-timeness of data processing.

Method used

By semantic jointing of the target user's historical search records, the joint coefficients and web page information of the joint semantics are obtained, and the model parameters are updated in combination with real-time web page information to improve the accuracy and real-timeness of search results.

Benefits of technology

It improves the search efficiency of search engines and the directional accuracy of output results, ensures that the output content of the model meets user needs, and updates parameters in real time to adapt to data changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234403A_ABST
    Figure CN120234403A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to an intelligent data processing method based on an artificial intelligence search engine technology, which comprises the following steps of: performing semantic union on any keyword input in a target model by a target user and a historical search record of the target user to obtain at least one joint semantic; according to the occurrence frequency of the historical keyword in each joint semantic in the historical search record, obtaining the joint coefficient of each joint semantic, and obtaining the target joint semantic of any keyword in combination with each piece of historical webpage information containing the joint semantic in the historical search record; and according to the difference between the historical webpage information and the real-time webpage information of each target joint semantic, updating parameters in the target model to obtain an updated target model, and according to the updated target model, outputting a search result corresponding to the search content. And the accuracy of the output result of the large language model in the artificial intelligence search engine technology is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a data intelligent processing method based on artificial intelligence search engine technology. Background Art

[0002] An artificial intelligence search engine is a new generation of search system that combines information retrieval and large language models, such as DeepSeek, Bing Chat, Wolfram Alpha, etc. It uses technologies such as natural language processing (semantic understanding, such as distinguishing between "Apple Inc." and "fruit apple"), large language models (such as GPT-4, Gemini, Claude, etc., for generating answers and summarizing information), and knowledge graphs (structuring and associating data such as people, events, locations, etc. to provide direct answers) to improve the intelligence, accuracy, and interaction experience of searches: users input search content into a pre-trained large language model, the natural language processing technology generates search keywords or query statements based on the search content, the search engine performs information retrieval according to the search keywords or query statements, and returns the retrieved results to the large language model. After integrating and analyzing the received results, the large language model finally outputs the results.

[0003] For large language models, some parameters in traditional large language models need to be predefined manually. However, with the change of data scenarios, the manually predefined parameters cannot guarantee the accuracy of the data processing results of the large language model. For example, in a financial evaluation model, with the continuous changes of factors such as the market environment and policies and regulations, traditional financial evaluation models need to be adjusted manually continuously, resulting in a high maintenance cost of the financial evaluation model and it is difficult to guarantee the accuracy of the data processing of the financial evaluation model. At the same time, in terms of information retrieval, traditional search engines mainly rely on keyword matching and cannot rely on historical search records, resulting in poor semantic understanding ability and inability to output the results required by users, that is, the accuracy of the output results is poor, which in turn leads to a poor user experience. For example, when a user inputs "apple", the search engine cannot accurately determine whether the user wants to know information about the fruit "apple" or the product dynamics of the technology company "Apple", resulting in the search engine returning a large amount of content that is not relevant to the user's actual needs to the large language model, resulting in poor directional accuracy of the output content of the large language model, and the user needs to input a qualifier again to obtain valuable information.

[0004] Therefore, how to improve the accuracy of the output results of the large language model in artificial intelligence search engine technology has become an urgent problem to be solved. Summary of the Invention

[0005] In view of this, an embodiment of the present invention provides a data intelligent processing method based on artificial intelligence search engine technology to solve the problem of how to improve the accuracy of the output results of large language models in artificial intelligence search engine technology.

[0006] An embodiment of the present invention provides a data intelligent processing method based on artificial intelligence search engine technology, and the method includes the following steps: According to the search content input by the target user in the target model at the current moment, at least one keyword is extracted. For any keyword, according to the historical search records of the target user in the target model, semantic association is performed on the any keyword to obtain at least one associated semantics of the any keyword. The historical search records include historical keywords corresponding to the historical search content of the target user, and historical web page information corresponding to each historical keyword; According to the occurrence frequency of the historical keywords in each of the associated semantics in the historical search records, the association coefficient of each of the associated semantics is obtained. According to each historical web page information containing the associated semantics in the historical search records and the association coefficient of each of the associated semantics, the target associated semantics of the any keyword is obtained; According to the target associated semantics of each keyword, the real-time web page information of each target associated semantics is obtained. According to the difference between the historical web page information and the real-time web page information of each target associated semantics, the parameters in the target model are updated to obtain an updated target model. According to the updated target model, the search results corresponding to the search content are output.

[0007] Preferably, the obtaining the association coefficient of each of the associated semantics according to the occurrence frequency of the historical keywords in each of the associated semantics in the historical search records includes: For any associated semantics, in the historical search records, at least one target search record of the historical keywords in the any associated semantics and the time stamp of each target search record are obtained. The time interval between the time stamp of each target search record and the current moment is calculated, and the reciprocal of each time interval is linearly normalized to obtain the time feature value of each target search record; Calculate the proportion of the number of all target search records in the number of all historical search records to obtain the search frequency of the historical keywords in the any associated semantics; Perform weighted summation on the search frequency and the average value of the time feature values of all target search records to obtain the association coefficient of the any associated semantics.

[0008] Preferably, obtaining the target combined semantics of any keyword according to each piece of historical web page information containing the combined semantics in the historical search record and the combination coefficient of each combined semantics includes: For any combined semantics, in the historical search record, obtain the number of all pieces of historical web page information containing the any combined semantics to obtain the number of web page information of the any combined semantics, obtain the number of web page information of all combined semantics of the any keyword, and calculate the proportion of the number of web page information of the any combined semantics in the number of web page information of all combined semantics of the any keyword to obtain the web page frequency of the any combined semantics; Perform linear normalization on the web page frequency of the any combined semantics to obtain a normalized web page frequency value, and calculate the product between the normalized web page frequency value of the any combined semantics and the combination coefficient to obtain the expected possible probability of the any combined semantics; Obtain the expected possible probabilities of all combined semantics of the any keyword, and obtain the target combined semantics of the any keyword according to the expected possible probabilities of all combined semantics of the any keyword.

[0009] Preferably, obtaining the target combined semantics of any keyword according to the expected possible probabilities of all combined semantics of the any keyword includes: Among the expected possible probabilities of all combined semantics of the any keyword, select the combined semantics corresponding to the maximum value as the target combined semantics of the any keyword.

[0010] Preferably, updating the parameters in the target model according to the difference between the historical web page information and the real-time web page information of each target combined semantics to obtain an updated target model includes: For any real-time web page information, in the historical search record, obtain the target historical web page information with the same website address as the any real-time web page information at different timestamps, and form the any real-time web page information and the target historical web page information into the web page information to be analyzed; According to the timestamp corresponding to each piece of web page information to be analyzed, perform text comparison on two pieces of web page information to be analyzed corresponding to every two adjacent timestamps to obtain a comparison file of two pieces of web page information to be analyzed corresponding to every two adjacent timestamps, and obtain the information difference degree of each comparison file according to the distribution of the difference data in each comparison file; According to the information difference degrees of all comparison files, determine whether the any real-time web page information has changed data. If the any real-time web page information has changed data, record the any real-time web page information as the updated data source web page information of the target model; Obtain the information of all web pages of the updated data sources. According to the information of all web pages of the updated data sources, update the parameters in the target model to obtain the updated target model.

[0011] Preferably, obtaining the information difference degree of each of the comparison documents according to the distribution of the difference data in each of the comparison documents includes: For any one of the comparison documents, if there is no difference data in the any one of the comparison documents, set the information difference degree of the any one of the comparison documents to 0; If there is difference data in the any one of the comparison documents, subtract the reciprocal of the number of all difference data in the any one of the comparison documents from the constant 1 to obtain the first difference degree of the any one of the comparison documents; In the any one of the comparison documents, obtain the number of characters between every two adjacent difference data, and subtract the reciprocal of the average value of all the number of characters from the constant 1 to obtain the second difference degree of the any one of the comparison documents; Calculate the average between the first difference degree and the second difference degree to obtain the information difference degree of the any one of the comparison documents.

[0012] Preferably, judging whether the data of any real-time web page information has changed according to the information difference degrees of all the comparison documents includes: Among all the information difference degrees, obtain the occurrence times of each of the information difference degrees, and use the information difference degree corresponding to the maximum occurrence times as the information difference degree threshold; If the information difference degree of the comparison document corresponding to any real-time web page information is greater than the information difference degree threshold, determine that the data of the any real-time web page information has changed; If the information difference degree of the comparison document corresponding to any real-time web page information is less than or equal to the information difference degree threshold, determine that the data of the any real-time web page information has not changed.

[0013] Preferably, after obtaining the information of all web pages of the updated data sources, it further includes: If there is no information of the web page of the updated data source, output the search result corresponding to the search content according to the target model.

[0014] The beneficial effects of the embodiments of the present invention compared with the prior art are: The present invention extracts at least one keyword according to the search content input by a target user in a target model at the current moment. For any keyword, semantic combination is performed on the any keyword according to the historical search records of the target user in the target model, and at least one combined semantics of the any keyword is obtained. The historical search records include historical keywords corresponding to the historical search content of the target user and historical web page information corresponding to each historical keyword. According to the occurrence frequency of the historical keywords in each combined semantics in the historical search records, a combined coefficient of each combined semantics is obtained. According to each historical web page information containing the combined semantics in the historical search records and the combined coefficient of each combined semantics, the target combined semantics of the any keyword is obtained. According to the target combined semantics of each keyword, real-time web page information of each target combined semantics is obtained. According to the difference between the historical web page information and the real-time web page information of each target combined semantics, the parameters in the target model are updated to obtain an updated target model. According to the updated target model, a search result corresponding to the search content is output. Among them, by combining the historical search records of the target user, semantic combination is performed on the content searched by the target user at the current moment, and more user-demand-compliant target combined semantics are obtained, so as to improve the search efficiency of the search engine, and further improve the directional accuracy of the content output by the target model. At the same time, according to the real-time web page information of each target combined semantics, the parameters in the target model are updated in real time, so as to improve the accuracy and real-time nature of the output result of the target model. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 It is a flowchart of a data intelligent processing method for an artificial intelligence search engine technology provided in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The following details the embodiments of the present disclosure, and the examples of the embodiments are shown in the drawings. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present disclosure and should not be construed as a limitation to the present disclosure.

[0018] It should be noted that the terms "first", "second", etc. in the specification of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure.

[0019] To illustrate the technical solution of the present invention, it will be described below through specific embodiments.

[0020] See Figure 1 , which is a method flowchart of a data intelligent processing method based on artificial intelligence search engine technology provided by Embodiment 1 of the present invention. As Figure 1 shown, the method may include: Step S101, according to the search content input by the target user in the target model at the current moment, extract at least one keyword. For any keyword, according to the historical search records of the target user in the target model, perform semantic association on the any keyword to obtain at least one associated semantics of the any keyword. The historical search records include historical keywords corresponding to the historical search content of the target user, and historical web page information corresponding to each historical keyword.

[0021] Use the Jieba library in Python to segment the search content input by the target user in the target model at the current moment to obtain at least one keyword. Among them, the target model is a model constructed according to the content in a fixed web page, such as a financial risk assessment model, which is constructed through financial data such as tax rates and exchange rates in a fixed official web page to perform financial risk assessment according to the input content of the target user. The Jieba library is a prior art and will not be elaborated here. After obtaining the keywords, each keyword is retrieved through a search engine. Considering the situation where the semantics of keywords are not unique, for example, when the user inputs "apple", the search engine cannot accurately determine whether the user wants to know information related to the fruit "apple", resulting in the search engine possibly returning a large amount of data irrelevant to the needs of the target user to the target model, making the output result of the target model inaccurate.

[0022] Therefore, in the embodiments of the present invention, before retrieving each keyword through a search engine, according to the historical search records of the target user in the target model (including the historical keywords corresponding to the historical search content of the target user and the historical web page information corresponding to each historical keyword), the ColBERT model is used to perform semantic combination on each keyword to obtain at least one combined semantics for each keyword. Taking the kth keyword as an example, if the kth keyword is "apple", and there are historical keywords such as "model" and "memory" in the historical search records of the target user in the target model, then the combined semantics of "apple" are "apple model", "apple memory", etc., so as to analyze each combined semantics, preliminarily determine the semantics required by the user, improve the search efficiency of the search engine, and further improve the directional accuracy of the content output by the target model. Among them, the ColBERT model is a prior art and will not be elaborated here.

[0023] Step S102: Obtain the combined coefficient of each combined semantics according to the occurrence frequency of the historical keywords in each combined semantics in the historical search records, and obtain the target combined semantics of any keyword according to each historical web page information containing the combined semantics in the historical search records and the combined coefficient of each combined semantics.

[0024] In order to improve the accuracy of judging the semantics required by the user, in the embodiments of the present invention, the combined coefficient of each combined semantics of the kth keyword is obtained according to the occurrence frequency of the historical keywords in each combined semantics of the kth keyword in the historical search records, which is used to characterize the possibility that the ith combined semantics is the semantics required by the user. Taking the ith combined semantics of the kth keyword as an example, if the occurrence frequency of the historical keyword in the ith combined semantics in the historical search records is relatively high, then the possibility that the ith combined semantics is the semantics required by the user is relatively large. The specific method for obtaining the combined coefficient of the ith combined semantics is as follows: In the historical search records, obtain at least one target search record of the historical keyword in the ith combined semantics and the time stamp of each target search record, calculate the time interval between the time stamp of each target search record and the current moment, and perform linear normalization on the reciprocal of each time interval to obtain the time feature value of each target search record. Among them, linear normalization is a prior art and will not be elaborated here; Calculate the proportion of the number of all target search records in the number of all historical search records to obtain the search frequency of the historical keyword in the ith combined semantics; Perform weighted summation on the search frequency and the average value of the time feature values of all target search records to obtain the combined coefficient of the ith combined semantics.

[0025] In one embodiment, the formula for calculating the combination coefficient of the i-th combined semantics is as follows:

[0026] Wherein, represents the combination coefficient of the i-th combined semantics, represents the first weight, n represents the number of all target search records, represents the time interval between the j-th target search record and the current moment, represents the second weight, represents the search frequency of the historical keywords in the i-th combined semantics, and norm() represents the linear normalization function.

[0027] It should be noted that, the smaller, the larger, indicating that the target user searches for the historical keywords in the i-th combined semantics multiple times in a short period of time, and thus the larger, the greater the possibility that the i-th combined semantics is the semantics required by the user. Since there may be a situation where the target user searched for the historical keywords in the i-th combined semantics more frequently in the early stage, the search frequency of the historical keywords in the i-th combined semantics is used as the main reference, and is set, There is no limitation here, and the implementer can set it according to the specific scenario.

[0028] Since semantic combination is an irregular combination, there may be incorrect combined semantics. If only analyzing based on the historical keywords in the i-th combined semantics, it is impossible to determine whether the i-th combined semantics itself is correct. Therefore, it is necessary to determine whether the i-th combined semantics is correct according to its occurrence in the actual data. In the embodiments of the present invention, the i-th combined semantics is searched in all historical web page information in the historical search records of the target user through the Whoosh library, and the number of historical web page information containing the i-th combined semantics is obtained, which is used to characterize the occurrence of the i-th combined semantics in the actual data. Then, in combination with the combination coefficient of the i-th combined semantics, the expected possible probability of the i-th combined semantics is obtained, which is used to characterize the possibility that the i-th combined semantics is the semantics required by the user. Among them, the Whoosh library is a prior art and will not be elaborated here. The specific method for obtaining the expected possible probability of the i-th combined semantics is as follows: In the historical search records, obtain the number of all historical web page information containing the i-th combined semantics to get the number of web page information of the i-th combined semantics, obtain the number of web page information of all combined semantics of the k-th keyword, calculate the proportion of the number of web page information of the i-th combined semantics in the number of web page information of all combined semantics of the k-th keyword to obtain the web page frequency of the i-th combined semantics; Linearly normalize the web page frequency of the i-th combined semantics to obtain the normalized web page frequency value, and calculate the product between the normalized web page frequency value of the i-th combined semantics and the combination coefficient to obtain the expected possible probability of any of the combined semantics.

[0029] In one embodiment, the calculation formula for the expected possible probability of the i-th combined semantics is:

[0030] Wherein, represents the expected possible probability of the i-th combined semantics, represents the combination coefficient of the i-th combined semantics, represents the web page frequency of the i-th combined semantics, and norm() represents the linear normalization function.

[0031] It should be noted that, the larger the , the more times the i-th combined semantics appears in the actual data, and thus the larger the , the greater the possibility that the i-th combined semantics is the semantics required by the user; if is small, it indicates that the i-th combined semantics is an incorrect combined semantics, that is , the i-th combined semantics must not be the semantics required by the user; the larger the , it indicates that the target user has searched for the historical keywords in the i-th combined semantics multiple times in a short period of time, and thus the larger the , the greater the possibility that the i-th combined semantics is the semantics required by the user.

[0032] Similarly, obtain the expected possible probabilities of all combined semantics of the k-th keyword input by the target user in the target model, and select the combined semantics corresponding to the maximum value among the expected possible probabilities of all combined semantics as the target combined semantics of the k-th keyword to represent the semantics required by the user.

[0033] Step S103, according to the target combined semantics of each keyword, obtain the real-time web page information of each target combined semantics, and update the parameters in the target model according to the difference between the historical web page information and the real-time web page information of each target combined semantics to obtain the updated target model, and output the search result corresponding to the search content according to the updated target model.

[0034] According to the obtaining method of the target combined semantics of the k-th keyword in step S102, obtain the target combined semantics of each keyword input by the target user in the target model. Retrieve the real-time web page information of each target combined semantics through a search engine, and the target model analyzes and integrates the information content in the real-time web page information of each target combined semantics, and outputs the analyzed and integrated result to the target user.

[0035] However, considering that some models themselves are not real-time, such as financial risk assessment models, some of the parameters in the financial risk assessment model are obtained based on the data in some official fixed web pages and a well-known formula. With the adjustment of influencing factors such as economic prices and exchange rates, the data in these fixed web pages will change, and the parameters in the original financial risk assessment model may not be applicable to the current financial risk assessment, resulting in inaccurate output results of the original financial risk assessment model. Since the target model will save the web page information to the historical search record of the target user in the target model every time it uses a certain web page information, in the embodiment of the present invention, after obtaining each target joint semantic real-time web page information, first compare these real-time web page information with the historical web page information in the historical search record to determine whether the data in these real-time web page information has changed. If it has changed, the parameters of the target model need to be updated to improve the accuracy of the output results of the target model.

[0036] In the embodiment of the present invention, taking the hth real-time web page information as an example, in the historical search record of the target user in the target model, obtain the target historical web page information with the same web address as the hth real-time web page information at different timestamps, and form the to-be-analyzed web page information by combining the hth real-time web page information and the target historical web page information. According to the timestamp corresponding to each to-be-analyzed web page information, use WinMerge to perform text comparison on two to-be-analyzed web page information corresponding to two adjacent timestamps. There is no limitation here, and the implementer can set the text comparison tool according to the specific scenario to obtain the comparison file of two to-be-analyzed web page information corresponding to two adjacent timestamps. Since the parameters of the target model have a direct relationship with digital data, in the comparison file, the highlighted data is the digital data with differences between the two to-be-analyzed web page information, and the highlighted data is recorded as the difference data. Among them, using WinMerge for text comparison is a prior art and will not be elaborated here. Further, according to the distribution of the difference data in each comparison file, obtain the information difference degree of each comparison file. Specifically: For any comparison file, if there is no difference data in the any comparison file, set the information difference degree of the any comparison file to 0; If there is difference data in the any comparison file, subtract the reciprocal of the number of all difference data in the any comparison file from the constant 1 to obtain the first difference degree of the any comparison file; In the any comparison file, obtain the number of characters between every two adjacent difference data, and subtract the reciprocal of the average value of all character numbers from the constant 1 to obtain the second difference degree of the any comparison file; Calculate the average value between the first difference degree and the second difference degree to obtain the information difference degree of the any comparison file.

[0037] In one embodiment, taking the y-th comparative document as an example, the calculation formula for the information difference degree of the y-th comparative document is as follows:

[0038] Wherein, represents the information difference degree of the y-th comparative document, represents the number of all difference data in the y-th comparative document, represents the number of characters between the s-th pair of adjacent difference data in the y-th comparative document.

[0039] It should be noted that since there may be data such as time that changes daily in the web page information, these data have nothing to do with the parameters of the target model, and these data may be concentrated together with a small positional gap between them. Therefore, the smaller, the fewer the number of characters between adjacent two difference data, that is, the smaller the distance between each adjacent two difference data, and thus the smaller, the less likely that the adjacent two difference data are parameters related to the target model; the smaller, the fewer the number of difference data in the y-th comparative document, the smaller the difference between the two web page information to be analyzed corresponding to the y-th comparative document, and thus the smaller, the less likely that the adjacent two difference data are parameters related to the target model.

[0040] Similarly, obtain the information difference degrees of all comparative documents. Among all the information difference degrees, obtain the occurrence times of each information difference degree, and take the information difference degree corresponding to the maximum occurrence times as the information difference degree threshold. If the information difference degree of the comparative document corresponding to the h-th real-time web page information is greater than the information difference degree threshold, it is determined that the data of the h-th real-time web page information has changed; if the information difference degree of the comparative document corresponding to the h-th real-time web page information is less than or equal to the information difference degree threshold, it is determined that the data of the h-th real-time web page information has not changed.

[0041] If the data of the h-th real-time web page information changes, the h-th real-time web page information is recorded as the web page information of the update data source for the target model. Similarly, it is judged whether the data of all real-time web page information changes to obtain all web pages of the update data source, and the parameters in the target model are updated according to the information of all web pages of the update data source to obtain the updated target model. For example, an industry evaluation model is constructed. This model is based on the number of companies, company levels, the number of employees, the talent gap, the average salary, etc. in each industry, and finally obtains a comprehensive evaluation of a certain industry to provide reference for job seekers. Suppose the key built-in parameters of the industry evaluation model are: the number of companies; company levels; the number of employees; the talent gap; the average salary, etc. The target user inputs "communication" in this model, and after semantic combination of "communication", the target combined semantics is obtained as "communication software". The average salary of "communication software" in the current industry evaluation model is 10,000, but according to the information of the update data source web page, the current average salary of "communication software" is 8,000. Then the average salary of "communication software" in the industry evaluation model is updated to 8,000 to obtain the updated industry evaluation model.

[0042] According to the updated target model, analyze and integrate the information content in the real-time web page information of each target combined semantics, and output the analyzed result to the user.

[0043] Specifically, if there is no information of the update data source web page, it means that the parameters in the target model do not need to be updated. Then the target model analyzes and integrates the real-time web page information of each target combined semantics returned by the search engine, and outputs the analyzed and integrated result to the target user.

[0044] It is worth noting that the focus of the present invention is: how to perform semantic combination according to the search content of the target user to improve the search efficiency of the search engine, thereby improving the directional accuracy of the content output by the target model. At the same time, according to the real-time web page information retrieved by the search engine, it is judged whether the parameters in the target model need to be updated, and then the parameters in the target model are updated to improve the accuracy of the output result of the target model. Among them, the update of the parameters in the target model is the prior art and will not be elaborated here.

[0045] In summary, the present invention extracts at least one keyword according to the search content input by the target user in the target model at the current moment. For any keyword, semantic combination is performed on the any keyword according to the historical search records of the target user in the target model, and at least one combined semantics of the any keyword is obtained. The historical search records include historical keywords corresponding to the historical search content of the target user and historical web page information corresponding to each historical keyword; according to the occurrence frequency of the historical keywords in each combined semantics in the historical search records, the combined coefficient of each combined semantics is obtained, and according to each historical web page information containing the combined semantics in the historical search records and the combined coefficient of each combined semantics, the target combined semantics of the any keyword is obtained; according to the target combined semantics of each keyword, the real-time web page information of each target combined semantics is obtained, and according to the difference between the historical web page information and the real-time web page information of each target combined semantics, the parameters in the target model are updated to obtain an updated target model, and according to the updated target model, the search results corresponding to the search content are output. Among them, by combining the historical search records of the target user, semantic combination is performed on the content searched by the target user at the current moment, and more user-demand-compliant target combined semantics are obtained, improving the search efficiency of the search engine and further improving the directional accuracy of the content output by the target model; at the same time, according to the real-time web page information of each target combined semantics, the parameters in the target model are updated in real time, improving the accuracy and real-time nature of the output results of the target model.

[0046] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A data intelligent processing method based on artificial intelligence search engine technology, characterized in that, The data intelligent processing method based on artificial intelligence search engine technology includes: According to the search content input by the target user in the target model at the current moment, at least one keyword is extracted. For any keyword, semantic association is performed on the any keyword according to the historical search records of the target user in the target model, and at least one combined semantics of the any keyword is obtained. The historical search records include historical keywords corresponding to the historical search content of the target user, and historical web page information corresponding to each historical keyword; According to the occurrence frequency of the historical keywords in each combined semantics in the historical search records, the combined coefficient of each combined semantics is obtained. According to each historical web page information containing the combined semantics in the historical search records and the combined coefficient of each combined semantics, the target combined semantics of the any keyword is obtained; According to the target combined semantics of each keyword, the real-time web page information of each target combined semantics is obtained. According to the difference between the historical web page information and the real-time web page information of each target combined semantics, the parameters in the target model are updated to obtain an updated target model. According to the updated target model, the search results corresponding to the search content are output.

2. The data intelligent processing method based on artificial intelligence search engine technology according to claim 1, characterized in that, The obtaining of the combined coefficient of each combined semantics according to the occurrence frequency of the historical keywords in each combined semantics in the historical search records includes: For any combined semantics, in the historical search records, at least one target search record of the historical keywords in the any combined semantics and the time stamp of each target search record are obtained. The time interval between the time stamp of each target search record and the current moment is calculated, and the reciprocal of each time interval is linearly normalized to obtain the time feature value of each target search record; Calculate the proportion of the number of all target search records in the number of all historical search records to obtain the search frequency of the historical keywords in the any combined semantics; Perform weighted summation on the search frequency and the average value of the time feature values of all target search records to obtain the combined coefficient of the any combined semantics.

3. The data intelligent processing method based on artificial intelligence search engine technology according to claim 1, characterized in that The obtaining of the target combined semantics of the any keyword according to each historical web page information containing the combined semantics in the historical search records and the combined coefficient of each combined semantics includes: For any combined semantics, in the historical search records, the number of all historical web page information containing the any combined semantics is obtained to obtain the number of web page information of the any combined semantics. The number of web page information of all combined semantics of the any keyword is obtained. Calculate the proportion of the number of web page information of the any combined semantics in the number of web page information of all combined semantics of the any keyword to obtain the web page frequency of the any combined semantics; Linearly normalize the web page frequency of the any combined semantics to obtain a normalized web page frequency value. Calculate the product of the normalized web page frequency value of the any combined semantics and the combined coefficient to obtain the expected possible probability of the any combined semantics; Obtain the expected possible probabilities of all combined semantics of any one of the keywords, and obtain the target combined semantics of any one of the keywords according to the expected possible probabilities of all combined semantics of any one of the keywords.

4. The data intelligent processing method based on artificial intelligence search engine technology according to claim 3, wherein, The obtaining the target combined semantics of any one of the keywords according to the expected possible probabilities of all combined semantics of any one of the keywords includes: Among the expected possible probabilities of all combined semantics of any one of the keywords, select the combined semantics corresponding to the maximum value as the target combined semantics of any one of the keywords.

5. The data intelligent processing method based on artificial intelligence search engine technology according to claim 1, wherein, The updating the parameters in the target model according to the differences between the historical web page information and the real-time web page information of each target combined semantics to obtain the updated target model includes: For any real-time web page information, in the historical search records, obtain the target historical web page information with the same website address as the any real-time web page information at different timestamps, and form the to-be-analyzed web page information by combining the any real-time web page information and the target historical web page information; According to the timestamps corresponding to each to-be-analyzed web page information, perform text comparison on two to-be-analyzed web page information corresponding to every two adjacent timestamps to obtain the comparison files of two to-be-analyzed web page information corresponding to every two adjacent timestamps, and obtain the information difference degree of each comparison file according to the distribution of the difference data in each comparison file; According to the information difference degrees of all comparison files, determine whether the any real-time web page information has changed data. If the any real-time web page information has changed data, then record the any real-time web page information as the updated data source web page information of the target model; Obtain all updated data source web page information, and update the parameters in the target model according to all updated data source web page information to obtain the updated target model.

6. The data intelligent processing method based on artificial intelligence search engine technology according to claim 5, wherein, The obtaining the information difference degree of each comparison file according to the distribution of the difference data in each comparison file includes: For any comparison file, if there is no difference data in the any comparison file, set the information difference degree of the any comparison file to 0; If there is difference data in the any comparison file, then subtract the reciprocal of the number of all difference data in the any comparison file from the constant 1 to obtain the first difference degree of the any comparison file; In the any comparison file, obtain the number of characters between every two adjacent difference data, and subtract the reciprocal of the average value of all character numbers from the constant 1 to obtain the second difference degree of the any comparison file; Calculate the average of the first difference degree and the second difference degree to obtain the information difference degree of the any comparison file.

7. The data intelligent processing method based on artificial intelligence search engine technology according to claim 5, characterized in that, The determining whether the any real-time web page information has changed data according to the information difference degrees of all comparison files includes: In all information difference degrees, obtain the occurrence times of each information difference degree, and use the information difference degree corresponding to the maximum occurrence times as the information difference degree threshold; If the information difference degree of the comparison file corresponding to the any real-time web page information is greater than the information difference degree threshold, determine that the any real-time web page information has changed data; If the information difference degree of the comparison file corresponding to any of the real-time web page information is less than or equal to the information difference degree threshold, it is determined that no data change has occurred in any of the real-time web page information.

8. The data intelligent processing method based on artificial intelligence search engine technology according to claim 5, characterized in that After obtaining all the updated data source web page information, it further includes: If there is no updated data source web page information, the search result corresponding to the search content is output according to the target model.

Citation Information

Patent Citations

  • Website navigation service system based on artificial intelligence

    CN119249011A

  • Keyword-based search method and device, medium and equipment

    CN119557421A

  • Search result processing method, device, terminal, electronic device, and storage medium

    WO2020108608A1