Recommendation method and device, equipment and medium
By clustering and interactive data analysis of user history search texts and determining target clusters, the problem of inaccurate user needs identification in the existing recommendation system is solved, and the selection of recommended content that is more in line with user needs is achieved, and the recommendation effect is improved.
Patent Information
- Application Number
- CN202510353763.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
AI Technical Summary
The existing recommendation system fails to accurately identify the user's demand for specific topic content, resulting in poor recommendation results.
By clustering the user's historical search text, using interactive data to determine the target cluster, indicating the user's demand popularity based on the target cluster, thereby selecting candidate content for recommendation.
It improves the accuracy and efficiency of recommended content, meets users' needs for popular content, and improves the recommendation effect.
Smart Images

Figure CN120277268A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to the field of recommendation technology. Specifically, the present disclosure relates to a recommendation method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] Artificial intelligence is a discipline that studies how to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), including both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0003] With the development of computer technology and big data technology, behaviors such as consumption, entertainment, learning, and travel in people's lives are closely related to big data. In the operation of a business platform, it is usually necessary to actively recommend content to users to improve the corresponding business performance.
[0004] The methods described in this section are not necessarily methods that have been previously envisioned or adopted. Unless otherwise specified, no method described in this section should be considered prior art merely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention
[0005] The present disclosure provides a recommendation method, apparatus, electronic device, computer-readable storage medium, and computer program product.
[0006] According to one aspect of the present disclosure, there is provided a recommendation method, including: obtaining user characteristics of a target user, where the user characteristics are determined based on historical interaction content of the target user; determining recommended content from a plurality of candidate contents based on the user characteristics of the target user; and recommending the recommended content to the target user, where the plurality of candidate contents are determined by the following method: obtaining a plurality of historical search texts and interaction data corresponding to each historical search text; clustering the plurality of historical search texts to obtain a plurality of clustering clusters; determining at least one target clustering cluster based on the interaction data corresponding to the historical search texts included in each clustering cluster of the plurality of clustering clusters; and determining the plurality of candidate contents based on search results corresponding to the historical search texts included in the at least one target clustering cluster.
[0007] According to one aspect of the present disclosure, a recommendation device is provided, including: a first acquisition unit configured to acquire user characteristics of a target user, where the user characteristics are determined based on historical interaction content of the target user; a first determination unit configured to determine recommended content from a plurality of candidate contents based on the user characteristics of the target user; and a recommendation unit configured to recommend the recommended content to the target user, where the plurality of candidate contents are determined by a candidate content determination device, and the candidate content determination device includes: a second acquisition unit configured to acquire a plurality of historical search texts and interaction data corresponding to each historical search text; a clustering unit configured to cluster the plurality of historical search texts to obtain a plurality of clustering clusters; a second determination unit configured to determine at least one target clustering cluster based on the interaction data corresponding to the historical search texts included in each clustering cluster among the plurality of clustering clusters; and a third determination unit configured to determine the plurality of candidate contents based on search results corresponding to the historical search texts included in the at least one target clustering cluster.
[0008] According to one aspect of the present disclosure, an electronic device is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned recommendation method.
[0009] According to one aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to cause the computer to execute the above-mentioned recommendation method.
[0010] According to one aspect of the present disclosure, a computer program product is provided, including a computer program, where the computer program can implement the above-mentioned recommendation method when executed by a processor.
[0011] According to one or more embodiments of the present disclosure, it is possible to mine the user's demand for popular content categories based on the user's historical search content and interaction data, and then determine recommended content from candidate contents that meet the popular demand, thereby improving the recommendation accuracy and efficiency.
[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0013] The accompanying drawings exemplarily illustrate embodiments and form part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0014] Figure 1 FIG. shows a schematic diagram of an exemplary system in which various methods described herein can be implemented according to an exemplary embodiment of the present disclosure;
[0015] Figure 2 FIG. shows a flowchart of a recommendation method according to an exemplary embodiment of the present disclosure;
[0016] Figure 3 FIG. shows a schematic diagram of a clustering cluster update process according to an exemplary embodiment of the present disclosure;
[0017] Figure 4 FIG. shows a schematic diagram of an update process of a candidate content library according to an exemplary embodiment of the present disclosure;
[0018] Figure 5 FIG. shows a schematic diagram of a determination process of candidate content according to an exemplary embodiment of the present disclosure;
[0019] Figure 6 FIG. shows a schematic diagram of a recommendation process according to an exemplary embodiment of the present disclosure;
[0020] Figure 7 FIG. shows a structural block diagram of a recommendation device according to an exemplary embodiment of the present disclosure;
[0021] Figure 8 FIG. shows a structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure. Detailed Embodiments
[0022] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.
[0023] In this disclosure, unless otherwise specified, the use of terms such as "first" and "second" to describe various elements is not intended to limit the positional relationship, timing relationship, or relative importance of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, while in certain cases, based on the context description, they may also refer to different instances.
[0024] The terms used in the description of the various examples in this disclosure are for the purpose of describing specific examples only and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in this disclosure covers any one of the listed items and all possible combinations.
[0025] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0026] Figure 1 FIG. shows a schematic diagram of an exemplary system 100 in which the various methods and apparatuses described herein can be implemented according to embodiments of the present disclosure. Referring Figure 1 to, the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 can be configured to execute one or more applications.
[0027] In an embodiment of the present disclosure, the server 120 can run one or more services or software applications that enable the execution of a recommendation method.
[0028] In certain embodiments, the server 120 can also provide other services or software applications, which can include non-virtual environments and virtual environments. In certain embodiments, these services can be provided as web-based services or cloud services, for example, provided to users of the client devices 101, 102, 103, 104, 105, and / or 106 under a software as a service (SaaS) model.
[0029] In Figure 1In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or a combination thereof that may be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 may in turn utilize one or more client applications to interact with server 120 to utilize the services provided by these components. It should be understood that a variety of different system configurations are possible, which may differ from system 100. Thus, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.
[0030] Users may use client devices 101, 102, 103, 104, 105, and / or 106 to send historical search text and perform interactive operations so that server 120 or other components controlled by server 120 can collect the user's interaction data. The client device may provide an interface that enables the user of the client device to interact with the client device. The client device may also output information to the user via this interface. Although Figure 1 only six client devices are depicted, those skilled in the art will be able to understand that the present disclosure may support any number of client devices.
[0031] Client devices 101, 102, 103, 104, 105, and / or 106 may include various categories of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices, etc. These computer devices may run various categories and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT WindowsMobile OS, iOS, Windows Phone, Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, Internet-enabled gaming devices, etc. The client device is capable of executing various different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and may use various communication protocols.
[0032] Network 110 can be any type of network well-known to those skilled in the art, and it can support data communication using any one of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, token ring, wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (such as Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0033] Server 120 can include one or more general-purpose computers, dedicated server computers (such as PC (personal computer) servers, UNIX servers, midrange servers), blade servers, mainframes, server clusters, or any other suitable arrangement and / or combination. Server 120 can include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices of the server). In various embodiments, server 120 can run one or more services or software applications that provide the functions described below.
[0034] The computing units in server 120 can run one or more operating systems including any of the above operating systems as well as any commercially available server operating systems. Server 120 can also run any one of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0035] In some embodiments, server 120 can include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 can also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.
[0036] In some embodiments, the server 120 may be a server of a distributed system or a server integrated with a blockchain. The server 120 may also be a cloud server, or an intelligent cloud computing server or an intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, which is used to solve the defects of high management difficulty and weak business scalability existing in traditional physical hosts and virtual private server (VPS) services.
[0037] The system 100 may further include one or more databases 130. In certain embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. The databases 130 may reside in various locations. For example, the database used by the server 120 may be local to the server 120, or may be remote from the server 120 and may communicate with the server 120 via a network-based or dedicated connection. The databases 130 may be of different categories. In certain embodiments, the database used by the server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data to and from the database in response to commands.
[0038] In certain embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be databases of different categories, such as key-value repositories, object repositories, or conventional repositories supported by a file system.
[0039] Figure 1 The system 100 can be configured and operated in various ways so that the various methods and apparatuses described according to the present disclosure can be applied.
[0040] In the related art, the candidate content library for recommendation to users is usually customized by manual screening or directly determined based on popular search content, and fails to summarize the demand heat of users for content on specific topics (vertical categories).
[0041] Based on this, the present disclosure provides a recommendation method, which clusters based on the historical search texts of users, determines a target clustering cluster based on the interaction data corresponding to each historical search text, uses the target clustering cluster to indicate the demand heat of users for vertical content (such as content around a certain news, or hot categories such as education and medical care), and determines candidate recommended content based on this, so that the recommended content better meets the user's needs and improves the recommendation effect.
[0042] Figure 2 shows a flowchart of a recommendation method 200 according to an exemplary embodiment of the present disclosure. As Figure 2As shown, method 200 includes:
[0043] Step S210, obtaining user characteristics of a target user, where the user characteristics are determined based on historical interaction content of the target user;
[0044] Step S220, determining recommended content based on the user characteristics of the target user; and
[0045] Step S230, recommending the recommended content to the target user.
[0046] Among them, in step S220, determining recommended content from multiple candidate contents based on the user characteristics of the target user includes:
[0047] Step S221, obtaining multiple historical search texts and interaction data corresponding to each historical search text;
[0048] Step S222, clustering the multiple historical search texts to obtain multiple clustering clusters;
[0049] Step S223, determining at least one target clustering cluster based on the interaction data corresponding to the historical search texts included in each clustering cluster among the multiple clustering clusters;
[0050] Step S224, determining the multiple candidate contents based on the search results corresponding to the historical search texts included in the at least one target clustering cluster; and
[0051] Step S225, determining recommended content from the multiple candidate contents based on the user characteristics of the target user.
[0052] By applying the above method 200, semantic clustering is performed based on the user's historical search texts. Each clustering cluster includes one or more historical search texts related to a certain content category. In this case, screening can be performed based on the sum of the interaction data corresponding to the historical search texts included in each clustering cluster to obtain target clustering clusters with higher interaction heat. The target clustering clusters are used to indicate vertical content with higher user demand heat (such as content around a certain news, or content in hot categories such as education and medical care), and then candidate contents that can meet the user's popular demands are determined. By determining candidate recommended contents from the candidate contents obtained in the above manner, high-quality recommended contents that better meet the user's needs can be obtained, thereby improving the recommendation effect.
[0053] In some examples, the user characteristics in step S210 may be determined by mining the content that the target user is interested in based on the historical interaction content of the target user. In this example, the user characteristics may be, for example, content category labels indicating the content categories that the target user is interested in. In one example, the content that the target user is interested in may be filtered based on the interaction data of the target user for the historical interaction content. The interaction data may include, for example, information such as the number of accesses, access frequency, and browsing duration of the target user for the historical interaction content. The present disclosure does not limit the specific data form of the user characteristics as long as the user characteristics can indicate the preferred content of the target user.
[0054] According to some embodiments, the interaction data corresponding to each historical search text in step S221 includes at least one of the following: the sum of the number of times that multiple users search for the historical search text; the sum of the durations that the multiple users browse the search results corresponding to the historical search text; the sum of the number of times that the multiple users click on the search results corresponding to the historical search text; and the frequency of each user among the multiple users browsing the search results corresponding to the historical search text. Thereby, it can more accurately and comprehensively indicate the user's attention or revisit rate for specific content.
[0055] In the above example, the sum of the number of searches, the sum of the browsing durations, and the sum of the number of clicks of multiple users corresponding to each historical search text can indicate the user attention heat of the historical search text, while the browsing frequency of each user for the historical search text can indicate the user's revisit demand. By analyzing or filtering the historical search texts based on the data that can indicate the user's revisit demand, it is possible to more accurately determine the demand heat of the user for the corresponding content of the historical search text in the future time period.
[0056] In some examples, the browsing frequency of each user for the historical search text may be characterized by the following index: the average number of access days of the same user for the historical search text within a certain time period, so as to more accurately indicate the user's revisit demand for the search content.
[0057] According to some embodiments, clustering the multiple historical search texts in step S222 to obtain multiple clusters includes: inputting the multiple historical search texts into a first language model to obtain a text feature vector corresponding to each historical search text; and clustering based on the text feature vectors corresponding to each historical search text to obtain the multiple clusters. Thereby, it is possible to use the first language model to map the text to the semantic space to obtain the text feature vector, and realize simple and accurate clustering based on the text feature vector.
[0058] In some examples, the first language model can be pre-trained using a large-scale corpus. For example, it can be the ERNIE Bot large model, or it can also be a language model obtained by combining a semantic model and a graph structure model. As long as it can map historical search texts to the semantic space to obtain text feature vectors, the present disclosure does not limit the specific type and training method of the first language model.
[0059] In some examples, it can be to use the K-means clustering algorithm to cluster multiple text feature vectors to obtain multiple clustering clusters. In this case, the number of clustering clusters can be set manually in advance according to requirements, so that the clustering result can meet the content classification requirements for historical search texts.
[0060] In some examples, clustering multiple historical search texts to obtain multiple clustering clusters can also be achieved by other means. For example, it can be to determine the similarity between each historical search text based on various text similarity calculation methods (such as the longest common subsequence algorithm), and implement text clustering based on this.
[0061] According to some embodiments, the process of determining the multiple candidate contents further includes: recording the cluster centers of the multiple clustering clusters; in response to determining that there is a new search text, determining the updated clustering cluster to which the new search text belongs based on the similarity between the cluster centers of the multiple clustering clusters and the new search text; updating the at least one target clustering cluster based on the interaction data corresponding to the historical search texts included in each clustering cluster among the multiple clustering clusters and the interaction data corresponding to the new search text; and updating the multiple candidate contents based on the search results corresponding to the search texts included in the at least one target clustering cluster after the update. Thus, incremental update of the clustering clusters can be performed based on the new search text to ensure the timeliness of the candidate contents.
[0062] In some examples, incremental update of the clustering clusters can be performed regularly. For example, it can be to count the newly added new search texts based on a fixed time period, and perform batch update based on all the new search texts within this time period. In this example, the joining time of each new search text can be recorded, and then each new search text can be classified into the updated clustering cluster closest to this new search text (i.e., the clustering cluster with the highest similarity between the cluster center and this new search text) through the nearest neighbor matching method. In some examples, the new search texts can also be filtered based on certain conditions. For example, the new search texts that do not meet the preset conditions for the distances to multiple cluster centers can be excluded to avoid abnormal data points affecting the accuracy of the clustering information.
[0063] According to some embodiments, the process of determining the multiple candidate contents further includes: based on the new search text, re-determining the cluster center of the updated cluster; and for each historical search text included in the adjacent clusters of the updated cluster, based on the similarity between the re-determined cluster center of the updated cluster and the historical search text, re-determining the cluster to which the historical search text belongs. Thereby, local updates can be performed on multiple clusters based on new data points (new search texts), ensuring the accuracy of clustering information.
[0064] According to some embodiments, for each historical search text included in the adjacent clusters of the updated cluster, based on the similarity between the re-determined cluster center of the updated cluster and the historical search text, re-determining the cluster to which the historical search text belongs includes: in response to determining that the similarity between the cluster center of the updated cluster after update and the cluster center of the updated cluster before update is not greater than a first similarity threshold, re-determining the cluster to which the historical search text belongs. Thereby, local updates can be performed only when the degree of change in the cluster center before and after update is relatively large, restricting the number of update operations and saving computing resources.
[0065] In some examples, the similarity between the cluster center of the updated cluster after update and the cluster center of the updated cluster before update can be obtained by calculating the distance between the cluster center before update and the cluster center after update, so as to indicate the influence degree of the update operation based on the new data point on the cluster information, and further determine whether local updates need to be performed, saving computing resources while ensuring the accuracy of clustering information.
[0066] According to some embodiments, the process of determining the multiple candidate contents further includes: determining the number of update operations performed on the cluster center of the updated cluster, and wherein, in response to determining that the similarity between the cluster center of the updated cluster after update and the cluster center of the updated cluster before update is not greater than a first similarity threshold, re-determining the cluster to which the historical search text belongs includes: in response to determining that the similarity between the cluster center of the updated cluster after update and the cluster center of the updated cluster before update is not greater than a first similarity threshold, and in response to determining that the number of update operations is not greater than an operation number threshold, re-determining the cluster to which the historical search text belongs. Thereby, the loop number of local updates for the cluster can be restricted using the operation number threshold, saving computing resources.
[0067] According to some embodiments, re - determining the cluster center of the updated cluster based on the new search text includes: determining the weight of the original cluster center and the weight of the new search text of the updated cluster based on the number of historical search texts included in the updated cluster; and re - determining the cluster center of the updated cluster by performing a weighted calculation based on the weight of the original cluster center, the weight of the new search text, the original cluster center, and the new search text. Thus, the update process of the cluster center can be simplified by using the weighted average method, reducing the computational amount and improving the update efficiency.
[0068] According to some embodiments, in response to determining that there is a new search text, determining the updated cluster to which the new search text belongs based on the similarity between the cluster centers of the multiple clusters and the new search text includes: in response to determining that there is a new search text and in response to determining that there is no similar search text in the multiple clusters, determining the updated cluster to which the new search text belongs based on the similarity between the cluster centers of the multiple clusters and the new search text, where the similarity between the similar search text and the new search text is not less than a second similarity threshold. Method 200 further includes: in response to determining that there is the similar search text in the multiple clusters, updating the interaction data corresponding to the similar search text based on the interaction data corresponding to the new search text. By applying the above - mentioned technical means, when there is already a similar search text that meets the conditions in the clustering library, the interaction data corresponding to the new search text is directly incorporated into the similar search text, which can save computing resources.
[0069] Figure 3 FIG. shows a schematic diagram of the cluster update process according to an exemplary embodiment of the present disclosure. As Figure 3 shown, the cluster update process 300 includes the following steps:
[0070] Step S301, determine the cluster centers of the initial multiple clusters. In this example, the cluster centers of the initial multiple clusters indicate the information of the current multiple cluster centers.
[0071] Step S302, use the nearest - neighbor matching method to determine the cluster to which each of the multiple new data points belongs. In this example, the multiple new data points correspond to multiple new search texts, and each new search text can be classified into a cluster through the nearest - neighbor matching calculation.
[0072] Step S303, for each cluster, determine the updated cluster center of the cluster. By performing this step, the cluster center information of the updated cluster based on the new data can be obtained to ensure the accuracy of the cluster information affected by the new data points.
[0073] Step S304: Determine whether the change in the cluster center does not exceed a preset range, and determine whether the number of update operations exceeds the operation count threshold. By using the change range of the cluster center before and after the update and the operation count threshold of the update operation to limit the update operation, it is possible to limit the number of update operations while ensuring the accuracy of the clustering cluster information as much as possible, saving computing resources.
[0074] When the result of any condition judgment in step S304 is yes, this round of update operation can be ended, and the update results of multiple clustering clusters can be determined.
[0075] When the results of the condition judgments in step S304 are all no, steps S302 - S303 can be continued to implement iterative update and ensure the accuracy of the clustering cluster information.
[0076] Figure 4 Shows a schematic diagram of the update process of the candidate content library according to an exemplary embodiment of the present disclosure. As Figure 4 shown, the candidate content library includes a vectorization processing module 410, an initial search text set 420, an updated search text set 430, a clustering cluster information table 440, and a search text information table 450.
[0077] In this example, a first language model is deployed in the vectorization processing module 410, and thus the search text can be converted into a text feature vector. The initial search text set 420 and the updated search text set 430 are respectively used to store the user's historical search text and the new search text. The clustering cluster information table 440 is used to store clustering cluster information, such as the cluster centers of multiple clustering clusters, statistical information of the interaction data of the search texts included in each clustering cluster, etc. The search text information table 450 is used to store the mapping relationship between the search text and the clustering cluster, and can also store the statistical information of the interaction data of each search text.
[0078] In this example, the construction of the candidate content library may include the following steps:
[0079] Step S1: Batch request. The initial search text set 420 initiates a batch request to the vectorization processing module 410, specifically a request to batch convert multiple historical search texts into text feature vectors.
[0080] Step S2: Return vector. The vectorization processing module 410 returns the text feature vectors corresponding to each of the multiple historical search texts.
[0081] Step S3: Clustering. After obtaining the text feature vectors corresponding to each of the multiple historical search texts, clustering can be performed based on the multiple text feature vectors to obtain multiple clustering clusters, and then the information of the cluster centers of the multiple clustering clusters can be stored in the clustering cluster information table 440.
[0082] Step S4, iterative update. Based on the cluster centers of the determined multiple clusters, the mapping relationship between the search texts and the clusters in the search text information table 450 can be updated.
[0083] Step S5, query. When there is a new search text in the updated search text set 430, it is possible to query whether the new search text exists in the search text information table 450, and then determine whether it is necessary to determine the text feature vector based on the new search text and update the cluster information based on the query result.
[0084] Step S6, batch request. For new search texts that do not exist in the search text information table 450, the updated search text set 430 can initiate a batch request to the vectorization processing module 410, specifically a request to batch-convert multiple new search texts into text feature vectors.
[0085] Step S7, return vector. The vectorization processing module 410 returns the text feature vectors corresponding to each of the multiple new search texts.
[0086] Step S8, nearest neighbor matching. After obtaining the text feature vectors corresponding to each new search text, the nearest neighbor matching method can be used to calculate the similarity based on the text feature vectors corresponding to each new search text and the cluster centers stored in the cluster information table 440 to determine the cluster to which each text feature vector belongs.
[0087] Step S9, periodic update. By executing this step, the cluster information table 440 and the search text information table 450 can be periodically updated to ensure the accuracy of the cluster information and the search text information.
[0088] In some examples, the above cluster update process can be accelerated and optimized using parallel computing and distributed computing technologies. For example, the content to be updated can be distributed to multiple computing hardware, and the update process can be accelerated through parallel computing to improve the information update efficiency.
[0089] In some examples, in step S223, determining at least one target cluster based on the interaction data corresponding to the historical search texts included in each of the multiple clusters can be achieved in the following way: determining the statistical information of the interaction data corresponding to the historical search texts included in each cluster, for example, calculating the sum of the interaction data corresponding to the historical search texts included in each cluster, and then selecting the target cluster based on the statistical information of the interaction data to obtain the target cluster that can correspond to the popular interaction content of the user, and determining the candidate content that better meets the user's needs based on this.
[0090] In some examples, in step S224, determining the multiple candidate contents based on the search results corresponding to the historical search texts included in the at least one target clustering cluster can be implemented in the following manner: screening based on the interaction data corresponding to the historical search texts included in the target clustering cluster to obtain the contents that are most searched by users, and then using the contents that are most searched by users as candidate contents, and determining the recommended contents that better meet the user's needs therefrom.
[0091] In some examples, in step S224, determining the multiple candidate contents based on the search results corresponding to the historical search texts included in the at least one target clustering cluster can be implemented in the following manner: generating a customized search result page for the historical search texts included in the target clustering cluster, or screening the search results corresponding to the historical search texts included in the target clustering cluster to obtain high-quality search results corresponding to the historical search texts, and then being able to recommend the high-quality search results corresponding to the popular needs of the user to the user to improve the user experience.
[0092] In some examples, the multiple candidate contents can include multiple display types. For example, they can be the splash screen display content, floating window push content, search box recommended content, advertisement layer push content, etc. in the search result page, and then the recommendation quality of each content recommendation position can be comprehensively improved based on the high-quality candidate contents.
[0093] In some examples, it can be to configure a dedicated recall queue for the candidate content library including the above multiple candidate contents during the recommendation process to obtain high-quality recall results based on the above multiple candidate contents, so that the recommended contents better meet the user's needs and improve the recommendation accuracy.
[0094] According to some embodiments, the user characteristics include the content preference information of the target user. In step S225, determining the recommended content from the multiple candidate contents based on the user characteristics of the target user includes: determining the recommended content based on the relevance between the multiple candidate contents and the content preference information. Thus, accurate recommendation can be achieved by using the relevance between the user characteristics and the candidate contents, so that the recommended content better meets the user's needs.
[0095] In an actual application scenario, there is the following possibility: there is some high-quality search text in the full-volume historical search data of the search platform, and more comprehensive and accurate search results can be obtained based on the high-quality search text. In this case, the user can be assisted in optimizing the search results by recommending the high-quality search text to the user.
[0096] Based on this, according to some embodiments, the user characteristics include user search text. In step S225, determining recommended content from multiple candidate contents based on the user characteristics of the target user includes: determining the recommended content from the multiple candidate contents based on the relevance between the user search text and the multiple candidate contents. Thus, relevant recommended content can be determined based on the user search text in the user search scenario, assisting the user in optimizing the search results and improving the recommendation effect.
[0097] According to some embodiments, the process of determining the multiple candidate contents further includes: determining at least one target search text from the multiple historical search texts based on the interaction data corresponding to each historical search text; and determining the multiple candidate contents based on the at least one target search text. Thus, high-quality search texts can be filtered based on the interaction data corresponding to each historical search text to obtain search contents that better meet the requirements of the search scenario for recommendation, so that the recommended content better meets the user's needs.
[0098] In some examples, it can be based on metrics such as the number of access days, weekly access frequency, and number of access users corresponding to each historical search text to filter the search texts that can be used for recommendation, so as to improve the recommendation accuracy.
[0099] According to some embodiments, method 200 further includes: inputting the historical search texts and task description information included in the at least one target clustering cluster into a second language model to obtain rewritten search texts output by the second language model, where the task description information indicates the limiting conditions for the rewritten search texts, and the limiting conditions include at least one of a word count condition, a text content richness condition, and a text format condition, and the search results corresponding to the historical search texts included in the at least one target clustering cluster are obtained by searching based on the rewritten search texts. Thus, task description information can be constructed to rewrite the historical search texts, improving the quality of the search texts, and further optimizing the quality of the search results and recommended content.
[0100] In some examples, a text generation instruction for inputting into the second language model can be obtained based on the historical search texts and task description information. For example, the historical search texts and the limiting conditions for the rewritten search texts included in the task description information can be filled into an instruction template including text slots and limiting condition slots to obtain a rewritten search text generation instruction. By applying a preset template to determine the rewritten search text generation instruction, the convenience of text rewriting can be improved, and thus the recommendation efficiency can be improved.
[0101] In some examples, the second language model may be a generative language model pre-trained using a large-scale corpus. For example, it may be the Wenxin Yiyan large model. As long as it can achieve semantic rewriting of historical search texts, the present disclosure does not limit the specific type and training method of the second language model.
[0102] According to some embodiments, method 200 further includes: recording the most recent search time of the plurality of historical search texts; and for each historical search text in the plurality of historical search texts, in response to determining that the time interval between the most recent search time of the historical search text and the current time is not less than a time interval threshold, deleting the historical search text from the plurality of clustering clusters. Thereby, it is possible to delete search texts with too distant search times in the plurality of clustering clusters, reduce the data volume of the clustering library, and improve the computing efficiency while ensuring the timeliness of candidate content.
[0103] According to some embodiments, method 200 further includes: obtaining user interaction information of the target user for the recommended content; and updating the plurality of candidate contents based on the user interaction information. Thereby, it is possible to use the user interaction information to indicate the recommendation effect of the recommended content, optimize the quality of the candidate content through posterior feedback analysis, and thus obtain recommended content that better meets the user's needs, improving the recommendation accuracy.
[0104] In some examples, posterior analysis may be performed based on metrics such as the click-through rate and quality score of the recommended content, and then high-quality recommended content may be screened out, and the recommendation effect may be improved by continuing to recommend high-quality content to other users.
[0105] Figure 5 FIG. shows a schematic diagram of the process for determining candidate content according to an exemplary embodiment of the present disclosure. As Figure 5 shown, the process for determining candidate content may include the following steps:
[0106] Step S501, collect user interaction data for the search text. In this step, the user interaction data corresponding to each search text may be obtained, such as the number of searches, the number of search users, the user browsing duration, etc.
[0107] Step S502, data filtering and cleaning. In this step, data filtering rules may be set to optimize the data quality. For example, duplicate data, data with missing values, and obviously unreasonable data may be cleaned, etc., to obtain more accurate user interaction data.
[0108] Step S503, determine whether the interaction data corresponding to each search text meets a preset condition. In this step, screening rules for screening search texts based on the interaction data may be set. For example, search content that can indicate the high revisit demand of users may be screened based on metrics such as the search volume and browsing duration.
[0109] Step S504, determine the search text filtering result. By applying the above Step S503, the search text filtering result that can indicate the high revisit demand of the user can be obtained.
[0110] Step S505, rewrite the filtered search text. In this example, a language model can be used to rewrite the filtered search text to improve the quality of the search text, thereby optimizing the quality of the search results and recommended content.
[0111] Step S506, recommend to the user based on the rewritten search text.
[0112] After the operation of recommending to the user is implemented using Step S506, Step S501 can be iteratively executed to collect the interaction data of the user with the recommended content, and the content available for recommendation can be further filtered according to the posterior feedback of the user interaction to further improve the quality of the recommended content.
[0113] Figure 6 shows a schematic diagram of the recommendation process according to an exemplary embodiment of the present disclosure. As Figure 6 shown, the recommendation process applies two modules: a candidate content library 610 and a distribution scenario 620.
[0114] In this example, the candidate content in the candidate content library 610 can be determined based on the following two steps:
[0115] Step S611, filter high revisit search texts based on interaction data. By applying this step, the popular search content of the user and the high revisit search texts that the user periodically repeats can be mined to obtain candidate content that can meet the popular needs of the user and provide it to the distribution scenario 620.
[0116] Step S612, cluster and statistically analyze historical search texts. By applying this step, the demand popularity of the user for content on specific vertical topics can be mined to obtain the popular access categories of the user, and then candidate content that can meet the popular needs of the user can be obtained and provided to the distribution scenario 620.
[0117] In this example, the following two steps can be executed in the distribution scenario 620:
[0118] Step S621, determine the recommended content based on user characteristics. By applying this step, the recommended content that better meets the user needs can be recommended to the user based on user characteristics to improve the recommendation accuracy.
[0119] Step S622: Collect the interaction data of the user. By applying this step, the posterior feedback statistics of the recommended content can be realized based on the user interaction data, and the content interaction feedback information is fed back to the candidate content library 610, so that the candidate content library 610 can further optimize the content screening result based on the user's interaction feedback to obtain higher-quality recommended candidate content.
[0120] According to one aspect of the present disclosure, a recommendation device is further provided. Figure 7 The block diagram of the recommendation device 700 according to an exemplary embodiment of the present disclosure is shown. As Figure 7 shown, the device 700 includes:
[0121] A first acquisition unit 710, configured to acquire the user characteristics of the target user, where the user characteristics are determined based on the historical interaction content of the target user;
[0122] A first determination unit 720, configured to determine the recommended content from multiple candidate contents based on the user characteristics of the target user; and
[0123] A recommendation unit 730, configured to recommend the recommended content to the target user,
[0124] wherein the multiple candidate contents are determined by a candidate content determination device 740, and the candidate content determination device includes:
[0125] A second acquisition unit 741, configured to acquire multiple historical search texts and the interaction data corresponding to each historical search text;
[0126] A clustering unit 742, configured to cluster the multiple historical search texts to obtain multiple clustering clusters;
[0127] A second determination unit 743, configured to determine at least one target clustering cluster based on the interaction data corresponding to the historical search texts included in each clustering cluster among the multiple clustering clusters; and
[0128] A third determination unit 744, configured to determine the multiple candidate contents based on the search results corresponding to the historical search texts included in the at least one target clustering cluster.
[0129] According to some embodiments, the clustering unit 742 includes: an input subunit, configured to input the multiple historical search texts into a first language model to obtain a text feature vector corresponding to each historical search text; and a clustering subunit, configured to perform clustering based on the text feature vector corresponding to each historical search text to obtain the multiple clustering clusters.
[0130] According to some embodiments, the candidate content determining device 740 further includes: a first recording unit configured to record the cluster centers of the plurality of clusters; a fourth determining unit configured to, in response to determining that there is a new search text, determine an updated cluster to which the new search text belongs based on the similarity between the cluster centers of the plurality of clusters and the new search text; a first updating unit configured to update the at least one target cluster based on the interaction data corresponding to the historical search texts included in each of the plurality of clusters and the interaction data corresponding to the new search text; and a second updating unit configured to update the plurality of candidate contents based on the search results corresponding to the search texts included in the updated at least one target cluster.
[0131] According to some embodiments, the first updating unit includes: a first determining subunit configured to re-determine the cluster center of the updated cluster based on the new search text; and a second determining subunit configured to, for each historical search text included in the adjacent clusters of the updated cluster, re-determine the cluster to which the historical search text belongs based on the similarity between the re-determined cluster center of the updated cluster and the historical search text.
[0132] According to some embodiments, the second determining subunit is configured to: in response to determining that the similarity between the cluster center of the updated cluster after update and the cluster center of the updated cluster before update is not greater than a first similarity threshold, re-determine the cluster to which the historical search text belongs.
[0133] According to some embodiments, the candidate content determining device 740 further includes: a fifth determining unit configured to determine the number of update operations performed on the cluster center of the updated cluster, and wherein the second determining subunit is configured to: in response to determining that the similarity between the cluster center of the updated cluster after update and the cluster center of the updated cluster before update is not greater than a first similarity threshold, and in response to determining that the number of update operations is not greater than an operation number threshold, re-determine the cluster to which the historical search text belongs.
[0134] According to some embodiments, the first determining subunit includes: a first determining module configured to determine the weight of the original cluster center of the updated cluster and the weight of the new search text based on the number of historical search texts included in the updated cluster; and a second determining module configured to re-determine the cluster center of the updated cluster by performing a weighted calculation based on the weight of the original cluster center, the weight of the new search text, the original cluster center, and the new search text.
[0135] According to some embodiments, the fourth determination unit is configured to: in response to determining that there is a new search text and in response to determining that there is no similar search text in the multiple clustering clusters, determine an updated clustering cluster to which the new search text belongs based on the similarity between the cluster centers of the multiple clustering clusters and the new search text, where the similarity between the similar search text and the new search text is not less than a second similarity threshold, and the candidate content determination device 740 further includes: a third update unit configured to, in response to determining that there is the similar search text in the multiple clustering clusters, update the interaction data corresponding to the similar search text based on the interaction data corresponding to the new search text.
[0136] According to some embodiments, the interaction data corresponding to each historical search text includes at least one of the following: the sum of the number of times that a user searches for the historical search text; the sum of the durations that multiple users browse the search results corresponding to the historical search text; the sum of the number of times that multiple users click on the search results corresponding to the historical search text; and the frequency of each user among the multiple users browsing the search results corresponding to the historical search text.
[0137] According to some embodiments, the user feature includes the content preference information of the target user, and the first determination unit 720 is configured to: determine the recommended content based on the relevance between the multiple candidate contents and the content preference information.
[0138] According to some embodiments, the user feature includes the user search text, and the first determination unit 720 is configured to: determine the recommended content from the multiple candidate contents based on the relevance between the user search text and the multiple candidate contents.
[0139] According to some embodiments, the candidate content determination device 740 further includes: a sixth determination unit configured to determine at least one target search text from the multiple historical search texts based on the interaction data corresponding to each historical search text, and the third determination unit is configured to determine the multiple candidate contents based on the at least one target search text.
[0140] According to some embodiments, the device 700 further includes: a rewriting unit configured to input the historical search text and the task description information included in the at least one target clustering cluster into a second language model to obtain a rewritten search text output by the second language model, where the task description information indicates a restriction condition for the rewritten search text, and the restriction condition includes at least one of a word count condition, a text content richness condition, and a text format condition, and the search results corresponding to the historical search text included in the at least one target clustering cluster are obtained by searching based on the rewritten search text.
[0141] According to some embodiments, the apparatus 700 further includes: a second recording unit configured to record the most recent search time of the plurality of historical search texts; and a deletion unit configured to, for each of the plurality of historical search texts, in response to determining that the time interval between the most recent search time of the historical search text and the current time is not less than a time interval threshold, delete the historical search text from the plurality of clustering clusters.
[0142] According to some embodiments, the apparatus 700 further includes: a third obtaining unit configured to obtain user interaction information of the target user with respect to the recommended content; and a fourth updating unit configured to update the plurality of candidate contents based on the user interaction information.
[0143] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0144] According to an aspect of the present disclosure, there is also provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned recommendation method.
[0145] According to an aspect of the present disclosure, there is also provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the above-mentioned recommendation method.
[0146] According to an aspect of the present disclosure, there is also provided a computer program product including a computer program, wherein the computer program implements the above-mentioned recommendation method when executed by a processor.
[0147] Referring to Figure 8 , a block diagram of an electronic device 800 that can be a server or a client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0148] AsFigure 8 As shown, device 800 includes a computing unit 801 which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0149] Multiple components in device 800 are connected to the I / O interface 805, including: an input unit 806, an output unit 807, a storage unit 808, and a communication unit 809. The input unit 806 can be any type of device capable of inputting information into device 800. The input unit 806 can receive input numerical or character information, and generate key signal inputs related to user settings and / or function controls of the electronic device, and can include but are not limited to a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 807 can be any type of device capable of presenting information, and can include but are not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 808 can include but are not limited to a magnetic disk, an optical disk. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and can include but are not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a BluetoothTM device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0150] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the recommendation method. For example, in some embodiments, the recommendation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the recommendation method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the recommendation method by any other suitable means (e.g., by means of firmware).
[0151] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0152] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0153] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0154] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0155] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0156] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0157] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0158] Although embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by their equivalent elements. In addition, the steps can be executed in an order different from that described in the present disclosure. Further, various elements in the embodiments or examples can be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein can be replaced by equivalent elements that emerge after the present disclosure.
Claims
1. A recommendation method, comprising: Obtaining user characteristics of a target user, where the user characteristics are determined based on historical interaction content of the target user; Determining recommended content from multiple candidate contents based on the user characteristics of the target user; and Recommending the recommended content to the target user, wherein the multiple candidate contents are determined by the following method: Obtaining multiple historical search texts and interaction data corresponding to each historical search text; Clustering the multiple historical search texts to obtain multiple clustering clusters; Determining at least one target clustering cluster based on the interaction data corresponding to the historical search texts included in each clustering cluster among the multiple clustering clusters; and Determining the multiple candidate contents based on the search results corresponding to the historical search texts included in the at least one target clustering cluster.
2. The method according to claim 1, wherein The clustering of the multiple historical search texts to obtain multiple clustering clusters includes: Inputting the multiple historical search texts into a first language model to obtain a text feature vector corresponding to each historical search text; and Clustering based on the text feature vectors corresponding to each historical search text to obtain the multiple clustering clusters.
3. The method according to claim 1 or 2, wherein The process of determining the multiple candidate contents further includes: Recording the cluster centers of the multiple clustering clusters; In response to determining that there is a new search text, determining an updated clustering cluster to which the new search text belongs based on the similarity between the cluster centers of the multiple clustering clusters and the new search text; Updating the at least one target clustering cluster based on the interaction data corresponding to the historical search texts included in each clustering cluster among the multiple clustering clusters and the interaction data corresponding to the new search text; and Updating the multiple candidate contents based on the search results corresponding to the search texts included in the updated at least one target clustering cluster.
4. The method according to claim 3, wherein The process of determining the multiple candidate contents further includes: Redetermining the cluster center of the updated clustering cluster based on the new search text; and For each historical search text included in the adjacent clustering clusters of the updated clustering cluster, redetermining the clustering cluster to which the historical search text belongs based on the similarity between the redetermined cluster center of the updated clustering cluster and the historical search text.
5. The method according to claim 4, wherein, The redetermining the clustering cluster to which each historical search text included in the adjacent clustering clusters of the updated clustering cluster belongs based on the similarity between the redetermined cluster center of the updated clustering cluster and the historical search text includes: In response to determining that the similarity between the cluster center of the updated clustering cluster after update and the cluster center of the updated clustering cluster before update is not greater than a first similarity threshold, redetermining the clustering cluster to which the historical search text belongs.
6. The method according to claim 5, wherein The process of determining the multiple candidate contents further includes: Determining the number of update operations performed on the cluster center of the updated clustering cluster, and wherein the redetermining the clustering cluster to which the historical search text belongs in response to determining that the similarity between the cluster center of the updated clustering cluster after update and the cluster center of the updated clustering cluster before update is not greater than a first similarity threshold includes: In response to determining that the similarity between the cluster center of the updated cluster after the update and the cluster center of the updated cluster before the update is not greater than a first similarity threshold, and in response to determining that the number of update operations is not greater than an operation number threshold, re-determine the cluster to which the historical search text belongs.
7. The method according to any one of claims 4-6, wherein, The re-determining the cluster center of the updated cluster based on the new search text includes: Determining the weight of the original cluster center of the updated cluster and the weight of the new search text based on the number of historical search texts included in the updated cluster; and Re-determining the cluster center of the updated cluster by performing a weighted calculation based on the weight of the original cluster center, the weight of the new search text, the original cluster center, and the new search text.
8. The method according to any one of claims 3-7, wherein The determining the updated cluster to which the new search text belongs based on the similarity between the cluster centers of the multiple clusters and the new search text in response to determining that there is a new search text includes: In response to determining that there is a new search text and in response to determining that there is no similar search text in the multiple clusters, determining the updated cluster to which the new search text belongs based on the similarity between the cluster centers of the multiple clusters and the new search text, where the similarity between the similar search text and the new search text is not less than a second similarity threshold. The method further includes: in response to determining that there is the similar search text in the multiple clusters, updating the interaction data corresponding to the similar search text based on the interaction data corresponding to the new search text.
9. The method according to any one of claims 1-8, wherein, The interaction data corresponding to each historical search text includes at least one of the following: The sum of the number of times multiple users search for the historical search text; The sum of the durations for which the multiple users view the search results corresponding to the historical search text; The sum of the number of times the multiple users click on the search results corresponding to the historical search text; and The frequency of each user among the multiple users viewing the search results corresponding to the historical search text.
10. The method according to any one of claims 1-9, wherein, The user characteristics include the content preference information of the target user, and the determining the recommended content from multiple candidate contents based on the user characteristics of the target user includes: Determining the recommended content based on the relevance between the multiple candidate contents and the content preference information.
11. The method according to any one of claims 1-10, wherein the user characteristics include user search texts, and the determining the recommended content from multiple candidate contents based on the user characteristics of the target user includes: Determining the recommended content from the multiple candidate contents based on the relevance between the user search texts and the multiple candidate contents.
12. The method according to claim 11, wherein, The determining process of the multiple candidate contents further includes: Determining at least one target search text from the multiple historical search texts based on the interaction data corresponding to each historical search text; and Determining the multiple candidate contents based on the at least one target search text.
13. The method according to any one of claims 1-12, further includes: Input the historical search texts and task description information included in the at least one target clustering cluster into a second language model to obtain rewritten search texts output by the second language model, where the task description information indicates the restrictive conditions for the rewritten search texts, and the restrictive conditions include at least one of a word count condition, a text content richness condition, and a text format condition. Among them, the search results corresponding to the historical search texts included in the at least one target clustering cluster are obtained by searching based on the rewritten search texts.
14. The method according to any one of claims 1-13, further comprising: Recording the most recent search time of the plurality of historical search texts; And For each historical search text among the plurality of historical search texts, in response to determining that the time interval between the most recent search time of this historical search text and the current time is not less than a time interval threshold, deleting this historical search text from the plurality of clustering clusters.
15. The method according to any one of claims 1-14, further comprising: Obtaining user interaction information of the target user for the recommended content; And Updating the plurality of candidate contents based on the user interaction information.
16. A recommendation device, comprising: A first acquisition unit configured to acquire user characteristics of a target user, where the user characteristics are determined based on the historical interaction content of the target user; A first determination unit configured to determine recommended content from a plurality of candidate contents based on the user characteristics of the target user; and A recommendation unit configured to recommend the recommended content to the target user, where the plurality of candidate contents are determined by a candidate content determination device, and the candidate content determination device includes: A second acquisition unit configured to acquire a plurality of historical search texts and interaction data corresponding to each historical search text; A clustering unit configured to cluster the plurality of historical search texts to obtain a plurality of clustering clusters; A second determination unit configured to determine at least one target clustering cluster based on the interaction data corresponding to the historical search texts included in each clustering cluster among the plurality of clustering clusters; and A third determination unit configured to determine the plurality of candidate contents based on the search results corresponding to the historical search texts included in the at least one target clustering cluster.
17. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; Wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-15.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause a computer to execute the method according to any one of claims 1-15.
19. A computer program product includes a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-15.