Content pushing method and device, electronic equipment and storage medium

By using keywords and entities of candidate content for multi-dimensional relevance matching and weighting when pushing content on the Internet, the problem of insufficient content push quality and diversity in existing technologies is solved, and more efficient content push and browsing are achieved.

CN116992119BActive Publication Date: 2026-08-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2022-09-21
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing technologies for pushing content over the internet rely on rudimentary matching methods, resulting in insufficient quality and diversity of content delivery and impacting browsing efficiency.

Method used

By acquiring keywords and entities from candidate content, combining them into various matching benchmark data, performing relevance matching, and weighting the initial relevance, the target content is determined and pushed.

Benefits of technology

It improved the quality and diversity of content delivery and increased browsing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116992119B_ABST
    Figure CN116992119B_ABST
Patent Text Reader

Abstract

This application discloses a content push method, apparatus, electronic device, and storage medium. The content push method acquires multiple candidate contents to be pushed, extracts candidate keywords and candidate entities from the candidate contents, and combines candidate keywords and candidate entities into multiple matching benchmark data. This allows for a more comprehensive and multi-dimensional relevance matching between candidate contents and historical content. Subsequently, the initial relevance under the multiple matching benchmark data is weighted to obtain the target relevance between the candidate contents and historical content. Based on the target relevance, the target content is determined from the multiple candidate contents and pushed. This achieves the effect of determining the target content from multiple dimensions, effectively enriching the types of target content, improving the quality and diversity of content push, and increasing the efficiency of content browsing. It can be widely applied in cloud technology, computer technology, and other technical fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a content delivery method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the rapid development of internet technology, diverse and timely content can be quickly and conveniently accessed online. Currently, when browsing content related to a specific entity online, other content related to that entity is usually pushed to the browsing page. Related technologies require matching relevant content before pushing this content; however, current matching methods are often rudimentary, reducing the quality and diversity of content pushes and impacting browsing efficiency. Summary of the Invention

[0003] The following is an overview of the subject matter described in detail in this application. This overview is not intended to limit the scope of the claims.

[0004] This application provides a content push method, apparatus, electronic device, and storage medium, which can improve the quality and diversity of content push and increase the efficiency of browsing content.

[0005] On the one hand, embodiments of this application provide a content push method, including:

[0006] Obtain multiple candidate contents to be pushed, extract candidate keywords from the candidate contents, extract candidate entities from the candidate contents, and combine the candidate keywords and candidate entities into multiple matching benchmark data;

[0007] Obtain the historical content that has been pushed, and perform relevance matching between the candidate content and the historical content based on various matching benchmark data to obtain the initial relevance between the candidate content and the historical content under various matching benchmark data.

[0008] The initial relevance under various matching benchmark data is weighted to obtain the target relevance between the candidate content and the historical content;

[0009] Based on the target relevance, a target content is determined from a plurality of candidate contents, and the target content is pushed. On the other hand, embodiments of this application also provide a content push device, including:

[0010] The data extraction module is used to obtain multiple candidate contents to be pushed, extract candidate keywords from the candidate contents, extract candidate entities from the candidate contents, and combine the candidate keywords and candidate entities into multiple matching benchmark data.

[0011] The relevance matching module is used to obtain the historical content that has been pushed, and to perform relevance matching between the candidate content and the historical content based on various matching benchmark data to obtain the initial relevance between the candidate content and the historical content under various matching benchmark data.

[0012] The relevance calculation module is used to weight the initial relevance under various matching benchmark data to obtain the target relevance between the candidate content and the historical content;

[0013] The push module is used to determine the target content from multiple candidate contents based on the target relevance and push the target content.

[0014] Furthermore, the aforementioned data extraction module is specifically used for:

[0015] A word network is constructed based on the adjacency relationships between candidate words in the candidate content, and a first key score for the candidate words is determined based on the word network.

[0016] The first occurrence frequency of the candidate words in the candidate content is counted, and the second key score of the candidate words is determined based on the first occurrence frequency;

[0017] The first key score and the second key score are weighted to obtain the target key score of the candidate words, and candidate keywords are determined from the candidate words based on the target key score.

[0018] Furthermore, the aforementioned data extraction module is specifically used for:

[0019] The words in the candidate content are filtered by part of speech to obtain multiple candidate words;

[0020] Multiple candidate words are slid through a sliding window of a preset width, and the candidate words located in the sliding window during the sliding process are connected to obtain a word network;

[0021] The first criticality score of the candidate word is determined based on the weight of the candidate word in the word network.

[0022] Furthermore, the aforementioned data extraction module is specifically used for:

[0023] The number of words in the candidate content is determined, and a first frequency score is obtained based on the quotient between the first occurrence frequency and the number of words.

[0024] Determine the first content quantity of all the candidate content and the second content quantity of the candidate content containing the candidate words, and obtain the second frequency score based on the quotient between the first content quantity and the second content quantity;

[0025] The second key score of the candidate word is obtained based on the first frequency score and the second frequency score.

[0026] Furthermore, the aforementioned data extraction module is specifically used for:

[0027] Extract the initial entity from the candidate content, and determine the second occurrence frequency of the initial entity in the reference content, wherein the reference content is content that is different from the category of the candidate content;

[0028] Based on the identification method corresponding to the second occurrence frequency, the initial entity is subjected to ambiguity identification to obtain the ambiguity identification result of the initial entity;

[0029] Based on the ambiguity recognition result, candidate entities for the candidate content are determined from the initial entities.

[0030] Furthermore, the aforementioned data extraction module is specifically used for:

[0031] When the frequency of the second occurrence is greater than the first frequency threshold, the initial entity is subjected to ambiguity identification based on the context information of the initial entity in the candidate content;

[0032] When the second occurrence frequency is greater than or equal to the second frequency threshold and the second occurrence frequency is less than the first frequency threshold, the initial entity is input into the pre-trained classification model to perform ambiguity recognition on the initial entity;

[0033] Wherein, the first frequency threshold is greater than the second frequency threshold.

[0034] Furthermore, the aforementioned data extraction module is specifically used for:

[0035] Based on the ambiguity identification result, the ambiguous entity is determined from the initial entity;

[0036] The initial entities other than the ambiguous entities are taken as entities to be determined, and the first entity score of the entities to be determined is determined according to the entity position of the entities to be determined in the candidate content.

[0037] Determine the number of paragraphs in the candidate content that contain the entity to be determined, and determine the second entity score of the entity to be determined based on the number of paragraphs.

[0038] The scores of the first entity and the second entity are weighted to obtain the target entity score of the entity to be determined, and the candidate entities of the candidate content are determined from the entities to be determined based on the target entity score.

[0039] Furthermore, the aforementioned relevance calculation module is specifically used for:

[0040] The matching weight corresponding to the matching benchmark data is determined based on the number of data types that make up the matching benchmark data.

[0041] The initial relevance under various matching benchmark data is weighted according to the corresponding matching weight to obtain the target relevance between the candidate content and the historical content.

[0042] Furthermore, the aforementioned relevance matching module is specifically used for:

[0043] Retrieve multiple pieces of content to be matched;

[0044] Extract the keywords or entities to be matched from the content to be matched, and classify the multiple contents to be matched according to the keywords or entities to be matched to obtain the content index of the multiple contents to be matched under different keywords or entities to be matched;

[0045] Extract historical keywords or historical entities from the historical content, and obtain multiple candidate contents to be pushed based on the content index corresponding to the historical keywords or historical entities.

[0046] Furthermore, the number of target contents is multiple, and the aforementioned push module is specifically used for:

[0047] Structural analysis is performed on the target content to obtain multiple structural features of the target content;

[0048] The initial structural score of the target content is determined based on each of the structural features.

[0049] The target structure score is obtained by weighting the scores of the multiple initial structures.

[0050] Based on the target structure score, multiple target contents are filtered, and the filtered target contents are pushed out.

[0051] Furthermore, the number of target contents is multiple, and the aforementioned push module is specifically used for:

[0052] Determine the title similarity and body text similarity between any two of the target contents;

[0053] The content similarity is obtained by weighting the title similarity and the body text similarity.

[0054] Multiple target contents are filtered based on the content similarity, and the filtered target contents are then pushed to the system.

[0055] On the other hand, embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described content push method.

[0056] On the other hand, embodiments of this application also provide a computer-readable storage medium storing a computer program, which is executed by a processor to implement the above-described content push method.

[0057] On the other hand, embodiments of this application also provide a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the content push method described above.

[0058] The embodiments of this application include at least the following beneficial effects: By acquiring multiple candidate contents to be pushed, extracting candidate keywords and candidate entities from the candidate contents, the matching basis between the candidate contents and historical content can be enriched based on the candidate keywords and candidate entities, improving the comprehensiveness and accuracy of the relevance matching between candidate contents and historical content. Furthermore, by combining candidate keywords and candidate entities into multiple matching benchmark data, the relevance matching between candidate contents and historical content can be performed more comprehensively from multiple perspectives, further improving the comprehensiveness and accuracy of the relevance matching between candidate contents and historical content. Subsequently, the initial relevance under multiple matching benchmark data is weighted to obtain the target relevance between candidate contents and historical content. Based on the target relevance, the target content is determined from multiple candidate contents and pushed. Since the target relevance is obtained by weighting the initial relevance under multiple matching benchmark data, the effect of determining the target content in multiple dimensions can be achieved, thereby effectively enriching the types of target content, improving the quality and diversity of content push, and improving the efficiency of content browsing.

[0059] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the application. Attached Figure Description

[0060] The accompanying drawings are used to provide a further understanding of the technical solutions of this application and constitute a part of the specification. They are used together with the embodiments of this application to explain the technical solutions of this application and do not constitute a limitation on the technical solutions of this application.

[0061] Figure 1 A schematic diagram illustrating an optional implementation environment provided for an embodiment of this application;

[0062] Figure 2 A schematic diagram of an optional flow of the content push method provided in an embodiment of this application;

[0063] Figure 3 This is an optional schematic diagram illustrating the relevance matching between candidate content and historical content, provided as an embodiment of this application.

[0064] Figure 4 A schematic diagram of an optional process for extracting candidate keywords provided in an embodiment of this application;

[0065] Figure 5 This is a schematic diagram of an optional process for sliding multiple candidate words provided in an embodiment of this application;

[0066] Figure 6 This is a schematic diagram of an optional structure of a word network provided in an embodiment of this application;

[0067] Figure 7 A schematic diagram of an optional process for extracting candidate entities provided in an embodiment of this application;

[0068] Figure 8 An optional structural diagram of the content index provided in the embodiments of this application;

[0069] Figure 9 A schematic diagram of an optional complete process of the content push method provided in the embodiments of this application;

[0070] Figure 10 A schematic diagram illustrating an optional partitioning of the matching strategy provided in an embodiment of this application;

[0071] Figure 11 An optional schematic diagram of the content push interface provided in the embodiments of this application;

[0072] Figure 12 Another optional schematic diagram of the content push interface provided in the embodiments of this application;

[0073] Figure 13 This is a schematic diagram of an optional structure of the content push device provided in the embodiments of this application;

[0074] Figure 14 This is a partial structural block diagram of a terminal provided in an embodiment of this application;

[0075] Figure 15 A partial structural block diagram of the server provided in an embodiment of this application. Detailed Implementation

[0076] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0077] It should be noted that in various specific embodiments of this application, when processing data related to the characteristics of the target object, such as target object attribute information or attribute information sets, is required, the permission or consent of the target object will be obtained first. Furthermore, the collection, use, and processing of this data will comply with the relevant laws, regulations, and standards of the relevant countries and regions. The target object can be a user. In addition, when embodiments of this application need to obtain target object attribute information, separate permission or consent from the target object will be obtained through pop-ups or redirection to a confirmation page. Only after obtaining the target object's separate permission or consent will the necessary target object-related data for the normal operation of the embodiments of this application be obtained.

[0078] To facilitate understanding of the technical solutions provided in the embodiments of this application, some key terms used in the embodiments of this application will be explained below:

[0079] Cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to achieve data computation, storage, processing, and sharing. Based on the cloud computing business model, cloud technology encompasses network technology, information technology, integration technology, management platform technology, and application technology. It can form resource pools, providing flexible and convenient on-demand access. Cloud computing technology will become a crucial support. Backend services of technical network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites. With the rapid development and application of the internet industry, every item may have its own identification mark in the future, requiring transmission to backend systems for logical processing. Data at different levels will be processed separately, and various industry data will require robust system support, which can only be achieved through cloud computing.

[0080] The solutions provided in this application relate to cloud computing, a cloud technology. Cloud computing refers to the delivery and usage model of IT infrastructure, meaning obtaining necessary resources through a network in an on-demand and easily scalable manner. In a broader sense, cloud computing refers to the delivery and usage model of services, meaning obtaining necessary services through a network in an on-demand and easily scalable manner. These services can be IT and software-related, internet-related, or other services. Cloud computing is a product of the development and integration of traditional computer and network technologies such as grid computing, distributed computing, parallel computing, utility computing, network storage technologies, virtualization, and load balancing. Driven by the development of the internet, real-time data streams, the diversification of connected devices, and the demands of search services, social networks, mobile commerce, and open collaboration, cloud computing has developed rapidly. Unlike previous parallel and distributed computing, the emergence of cloud computing will, conceptually, drive a revolutionary change in the entire internet model and enterprise management model.

[0081] Currently, when browsing content related to a certain entity via the Internet, other content related to that entity is usually pushed to the browsing page. In related technologies, when pushing other content related to the entity, it is necessary to first match the relevant content. However, the current matching methods are usually quite rudimentary.

[0082] For example, taking a security as an example, if the current push is content T1 corresponding to "Securities A", then the page containing this content will push content T2, content T3, ..., content Tn related to "Securities A". Among them, content T2, content T3, ..., content Tn are obtained by matching the text "Securities A", that is, content T2, content T3, ..., content Tn will contain "Securities A". Obviously, this matching method yields a single type of content and cannot comprehensively display the information about Securities A.

[0083] It is evident that the content pushed through the aforementioned methods provides limited information and suffers from severe content homogenization, thereby reducing the quality and diversity of content pushes and impacting the efficiency of content browsing.

[0084] Based on this, embodiments of this application provide a content push method, apparatus, electronic device, and storage medium, which can improve the quality and diversity of content push and increase the efficiency of browsing content.

[0085] Reference Figure 1 , Figure 1 This is a schematic diagram of an optional implementation environment provided in an embodiment of this application. The implementation environment includes a terminal 101 and a server 102, wherein the terminal 101 and the server 102 are connected through a communication network.

[0086] For example, server 102 can obtain multiple candidate contents to be pushed to terminal 101, extract candidate keywords from the candidate contents, extract candidate entities from the candidate contents, and combine the candidate keywords and candidate entities into multiple matching benchmark data; server 10 can obtain historical content that has been pushed to terminal 101, perform relevance matching between candidate content and historical content according to various matching benchmark data, and obtain the initial relevance between candidate content and historical content under various matching benchmark data; weight the initial relevance under multiple matching benchmark data to obtain the target relevance between candidate content and historical content; determine the target content from multiple candidate contents according to the target relevance, and push the target content to terminal 101 for display.

[0087] Server 102 acquires multiple candidate contents to be pushed, extracts candidate keywords and candidate entities from the candidate contents, and enriches the matching criteria between candidate contents and historical content based on the candidate keywords and candidate entities, thereby improving the accuracy and reliability of the relevance matching between candidate contents and historical content. Furthermore, by combining candidate keywords and candidate entities into multiple matching benchmark data, the relevance matching between candidate contents and historical content can be performed more comprehensively from multiple perspectives, further improving the accuracy and reliability of the relevance matching between candidate contents and historical content. Subsequently, the initial relevance under multiple matching benchmark data is weighted to obtain the target relevance between candidate contents and historical content. Based on the target relevance, the target content is determined from multiple candidate contents and pushed to terminal 101 for display. Since the target relevance is obtained by weighting the initial relevance under multiple matching benchmark data, the effect of determining the target content in multiple dimensions can be achieved, thereby effectively enriching the types of target content, improving the quality and diversity of content push, and improving the efficiency of content browsing on terminal 101.

[0088] Server 102 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Additionally, server 102 can also be a node server in a blockchain network.

[0089] Terminal 101 can be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, vehicle terminal, etc., but is not limited to these. Terminal 101 and server 102 can be directly or indirectly connected via wired or wireless communication, and this embodiment of the application does not impose any limitations.

[0090] The methods provided in this application can be applied to various technical fields, including but not limited to cloud technology, computer technology, and other technical fields.

[0091] Reference Figure 2 , Figure 2 This is an optional flowchart illustrating a content push method provided in an embodiment of this application. The content push method can be executed by a server, or by a terminal, or by a combination of a server and a terminal. The content push method includes, but is not limited to, the following steps 201 to 204.

[0092] Step 201: Obtain multiple candidate contents to be pushed, extract candidate keywords from the candidate contents, extract candidate entities from the candidate contents, and combine candidate keywords and candidate entities into multiple matching benchmark data.

[0093] Here, candidate content refers to the content to be pushed, that is, the content that the server is about to push to the terminal. Candidate keywords are the keywords in the candidate content. For example, assuming the candidate content is "Company X is preparing to hold a groundbreaking ceremony in region Y", then the candidate keywords could be "Company X", "region Y", and "groundbreaking ceremony". It is understood that the extracted candidate keywords can be determined according to the actual situation of the candidate content. The above example is only for illustrative purposes.

[0094] In this context, candidate entities refer to the entities within the candidate content. Entities can be objectively existing and distinguishable things or events. For example, candidate entities can be securities, industries, or events. For instance, a securities is the name of a listed company's stock; an industry is the industry corresponding to that listed company, such as the new energy industry or the chemical industry; and an event can be information such as the listed company's earnings release, new product launch, or company acquisition. Alternatively, candidate entities can be sports teams, regions, or events. It is understood that the extracted candidate entities can be determined based on the actual situation of the candidate content; the above examples are merely illustrative.

[0095] In one possible implementation, candidate keywords and candidate entities are combined into multiple matching benchmark data. This could be using candidate keywords as matching benchmark data, using candidate entities as matching benchmark data, or using both candidate keywords and candidate entities as matching benchmark data. In this case, there are three types of matching benchmark data.

[0096] In addition, candidate entities can be categorized, and different categories of candidate entities can be combined with candidate keywords to obtain multiple matching benchmark data. For example, candidate entities can be categorized into securities, industries, and events. When combining candidate keywords and candidate entities to form multiple matching benchmark data, it can be that candidate keywords, securities, and industries are used as matching benchmark data; candidate keywords and securities are used as matching benchmark data; candidate keywords and industries are used as matching benchmark data; and candidate keywords and events are used as matching benchmark data. In this case, the number of matching benchmark data is four.

[0097] In addition, candidate keywords can be categorized and combined with candidate entities according to their different levels to obtain multiple matching benchmark data. For example, candidate keywords can be divided into primary keywords and secondary keywords. When combining candidate keywords and candidate entities to form multiple matching benchmark data, one could use primary keywords, secondary keywords, and candidate entities as matching benchmark data; another could use primary keywords and candidate entities as matching benchmark data; and yet another could use secondary keywords and candidate entities as matching benchmark data. In this case, there are three types of matching benchmark data.

[0098] Among them, primary keywords are closer to the central idea of ​​the content than secondary keywords. Candidate keywords are graded, specifically by inputting them into a preset keyword grading model to obtain their corresponding levels. It is understood that the above example exemplifies the grading of candidate keywords into primary and secondary keywords; in practice, more levels can be defined based on actual circumstances, and this application does not impose limitations.

[0099] In addition, candidate keywords at different levels can be combined with candidate entities of different categories to obtain multiple matching benchmark data. For example, candidate keywords can be divided into primary keywords and secondary keywords, and candidate entities can be divided into securities, industry, and event. When combining candidate keywords and candidate entities into multiple matching benchmark data, the following combinations can be used: primary keywords, secondary keywords, and securities, industry, and event as matching benchmark data; primary keywords and securities, industry, and event as matching benchmark data; secondary keywords and securities, industry, and event as matching benchmark data; primary keywords, secondary keywords, and securities as matching benchmark data; primary keywords, secondary keywords, and industry as matching benchmark data; and primary keywords, secondary keywords, and event as matching benchmark data. In this case, the number of matching benchmark data is six.

[0100] It should be noted that the number of matching benchmark data in the above examples is only for illustrative purposes. In reality, the number of matching benchmark data can vary depending on the combination of different candidate keywords and candidate entities, and this application does not limit it.

[0101] In one possible implementation, since both candidate keywords and candidate entities are actually text, after extracting the candidate keywords and candidate entities, a deduplication operation can be performed on the candidate keywords and candidate entities to improve the efficiency of subsequent data processing.

[0102] Step 202: Obtain the historical content that has been pushed, and perform relevance matching between candidate content and historical content based on various matching benchmark data to obtain the initial relevance between candidate content and historical content under various matching benchmark data.

[0103] Historical content refers to content that has already been pushed to the terminal. For example, it could be content pushed by the server to the terminal that has already been viewed, or it could be content pushed by the server to the terminal that is currently being viewed. This application's embodiments do not impose any limitations. Historical content is generally different from candidate content.

[0104] When performing relevance matching between candidate content and historical content based on various matching benchmark data, the number of relevance matchings is consistent with the number of combinations of matching benchmark data.

[0105] For example, when candidate keywords are used as the matching benchmark data, candidate entities are used as the matching benchmark data, and both candidate keywords and candidate entities are used as the matching benchmark data, then the candidate content is matched against historical content based on the candidate keywords, the candidate content is matched against historical content based on the candidate entities, and the candidate content is matched against historical content based on both candidate keywords and candidate entities. Accordingly, the number of initial relevance scores obtained is consistent with the number of relevance matches.

[0106] In one possible implementation, relevance matching between candidate content and historical content is performed based on matching benchmark data. Specifically, historical keywords and entities can be extracted from historical content, and these historical keywords and entities can be combined according to the types of matching benchmark data for the candidate content to obtain the corresponding matching benchmark data for the historical content. When performing relevance matching between candidate content and historical content, the dimensions of the matching benchmark data for both candidate and historical content are the same. That is, when the matching benchmark data for candidate content is a candidate keyword, the matching benchmark data for historical content is a historical keyword; when the matching benchmark data for candidate content is a candidate entity, the matching benchmark data for historical content is a historical entity; and when the matching benchmark data for candidate content is both a candidate keyword and a candidate entity, the matching benchmark data for historical content is both historical keywords and historical entities.

[0107] For example, refer to Figure 3 , Figure 3This is an optional schematic diagram illustrating relevance matching between candidate content and historical content provided in an embodiment of this application. The candidate content includes keywords T1, T2, and T3, as well as entities S1, S2, and S3. The historical content includes keywords T2, T3, and T4, as well as entities S1, S4, and S5. When performing relevance matching between candidate content and historical content based on matching benchmark data, the first matching benchmark data for the candidate content can be keywords T1, T2, and T3; correspondingly, the first matching benchmark data for the historical content can be keywords T2, T3, and T4. The second matching benchmark data for the candidate content can be entities S1, S2, and S5. 3. Accordingly, the second matching benchmark data for historical content can be entities S1, S4, and S5; the third matching benchmark data for candidate content can be keywords T1, T2, T3, entities S1, S2, and S3; and the third matching benchmark data for historical content can be keywords T2, T3, T4, entities S1, S4, and S5. Then, the candidate content and the first matching benchmark of historical content are matched for relevance to obtain the first initial relevance R1; the candidate content and the second matching benchmark of historical content are matched for relevance to obtain the second initial relevance R2; and the candidate content and the third matching benchmark of historical content are matched for relevance to obtain the third initial relevance R3.

[0108] In one possible implementation, relevance matching between candidate content and historical content is performed based on the matching benchmark data. Specifically, the intersection of the matching benchmark data of candidate content and the matching benchmark data of historical content is calculated to obtain the intersection of matching data, and the union of the matching benchmark data of candidate content and the matching benchmark data of historical content is calculated to obtain the union of matching data. The initial relevance is obtained based on the quotient between the number of elements in the intersection of matching data and the number of elements in the union of matching data.

[0109] For example, to undertake Figure 3In the example shown, for the first matching baseline data, the intersection of the matching data is keyword T2 and keyword T3, and the union of the matching data is keyword T1, keyword T2, keyword T3, and keyword T4, so the first initial relevance R1 is 0.5; for the second matching baseline data, the intersection of the matching data is entity S1, and the union of the matching data is entity S1, entity S2, entity S3, entity S4, and entity S5, so the second initial relevance R2 is 0.2; for the third matching baseline data, the intersection of the matching data is keyword T2, keyword T3, and entity S1, and the union of the matching data is keyword T1, keyword T2, keyword T3, keyword T4, entity S1, entity S2, entity S3, entity S4, and entity S5, so the third initial relevance R3 is 0.33.

[0110] In one possible implementation, when matching candidate content with historical content based on the matching benchmark data, the matching benchmark data of the candidate content and the matching benchmark data of the historical content can be converted into their respective vectors. The initial relevance can then be obtained by calculating the similarity between the vectors corresponding to the matching benchmark data of the candidate content and the matching benchmark data of the historical content.

[0111] In one possible implementation, when performing relevance matching between candidate content and historical content based on various matching benchmark data, the historical content can be traversed to count the number of times the matching benchmark data of the candidate content is matched in the historical content. The initial relevance between the candidate content and the historical content is determined based on the number of matches. Specifically, a mapping relationship between the number of matches and the initial relevance can be preset. After determining the number of matches of the candidate content's matching benchmark data in the historical content, the corresponding initial relevance is determined according to the preset mapping relationship.

[0112] By acquiring multiple candidate content items to be pushed, extracting candidate keywords and candidate entities from the candidate content, the matching criteria between candidate content and historical content can be enriched based on the candidate keywords and candidate entities, thereby improving the comprehensiveness and accuracy of the relevance matching between candidate content and historical content. Furthermore, by combining candidate keywords and candidate entities into multiple matching benchmark data, the relevance matching between candidate content and historical content can be performed more comprehensively from multiple perspectives, further improving the comprehensiveness and accuracy of the relevance matching between candidate content and historical content.

[0113] Step 203: Weight the initial relevance under multiple matching benchmark data to obtain the target relevance between candidate content and historical content.

[0114] When weighting the initial relevance under multiple matching benchmark data, the matching weights corresponding to the matching benchmark data can be determined first.

[0115] For example, to undertake Figure 3 In the example shown, the matching weights corresponding to the first matching benchmark data, the second matching benchmark data, and the third matching benchmark data are a1, a2, and a3, respectively. Then the target relevance is a1*R1+a2*R2+a3*R3.

[0116] In one possible implementation, when weighting the initial relevance under multiple matching benchmark data to obtain the target relevance between candidate content and historical content, the matching weight corresponding to the matching benchmark data can be determined according to the number of data types that make up the matching benchmark data; and the initial relevance under multiple matching benchmark data is weighted according to the corresponding matching weight to obtain the target relevance between candidate content and historical content.

[0117] For example, when candidate keywords are used as the matching benchmark data, the number of data types is 1; when candidate keywords and candidate entities are used as the matching benchmark data, the number of data types is 2. More specifically, when there are multiple categories of candidate entities, such as candidate entities being divided into securities, industries, and events, the number of data types is 3 when candidate keywords, securities, and industries are used as the matching benchmark data.

[0118] In particular, the more data types there are, the greater the matching weight can be for the benchmark data, thus making the target relevance more reasonable.

[0119] In one possible implementation, when the number of data types in the matching benchmark data is the same, the matching weight can be further determined based on the data types that make up the matching benchmark data. Specifically, the priority of the data types that make up the matching benchmark data can be determined, and the matching weight corresponding to the matching benchmark data can be determined based on this priority.

[0120] For example, if the number of data types in the matching benchmark data is 2, the matching benchmark data can be a combination of keywords and industries, or a combination of keywords and events. In this case, based on the preset combination of data types, the priority of industries can be determined to be higher than that of events. Therefore, the matching weight of the matching benchmark data composed of keywords and industries will be greater than that of the matching benchmark data composed of keywords and events.

[0121] Step 204: Determine the target content from multiple candidate contents based on the target relevance, and push the target content.

[0122] In one possible implementation, the content can be sorted from high to low based on the relevance of the target. Based on the sorting result, candidate content ranked before the first ranking threshold is obtained as the target content. The first ranking threshold can be set according to the actual situation, such as 5 or 10, etc. This application embodiment does not limit it.

[0123] Based on this, after obtaining candidate content ranked before the first ranking threshold according to the ranking results, the candidate content ranked before the first ranking threshold can be further ranked according to the publication time. The candidate content ranked before the second ranking threshold is then selected as the target content. The second ranking threshold is less than the first ranking threshold. The second ranking threshold can be set based on the first ranking threshold; for example, when the first ranking threshold is 10, the second ranking threshold can be 5, etc. This embodiment does not impose such limitations. By further introducing a second ranking threshold to determine the target content, the precision of the target content determination can be effectively improved.

[0124] By weighting the initial relevance under multiple matching benchmark data, the target relevance between candidate content and historical content is obtained. Based on the target relevance, the target content is determined from multiple candidate contents and pushed. Since the target relevance is obtained by weighting the initial relevance under multiple matching benchmark data, the effect of determining the target content in multiple dimensions can be achieved. This can effectively enrich the types of target content, improve the quality and diversity of content push, and improve the efficiency of browsing content.

[0125] In one possible implementation, when extracting candidate keywords from candidate content, a word network can be constructed based on the adjacency relationships between candidate words in the candidate content. The first key score of the candidate words is determined based on the word network. The first occurrence frequency of the candidate words in the candidate content is counted, and the second key score of the candidate words is determined based on the first occurrence frequency. The first key score and the second key score are weighted to obtain the target key score of the candidate words. Candidate keywords are then determined from the candidate words based on the target key score.

[0126] In one possible implementation, a first key score and a second key score are introduced when extracting candidate keywords. The first key score is mainly obtained by constructing a word network. Specifically, candidate words are words in the candidate content. The candidate content can be segmented to obtain multiple candidate words. For example, if the candidate content is "Company X is preparing to hold a groundbreaking ceremony in region Y", then after segmenting the candidate content, the candidate words obtained can be "Company X", "preparing", "in", "region Y", "conducting", and "groundbreaking ceremony". Then, a word network can be constructed based on the adjacency relationships of these candidate words. Next, the candidate words, as nodes in the word network, can be used to iteratively calculate the score of each node based on the word network, thereby obtaining the first key score.

[0127] The second key score is primarily determined by the frequency of the first occurrence of the candidate word within the candidate content. For example, if the candidate content is "Company X is preparing to hold a groundbreaking ceremony in region Y, and all of Company X's executives will attend the ceremony," then the first occurrence frequency of the candidate word "Company X" is 2. The first occurrence frequencies of the other candidate words are determined similarly, and will not be elaborated further here. After determining the first occurrence frequency of the candidate words, when determining the second key score based on the first occurrence frequency, a pre-defined mapping relationship between the first occurrence frequency and the second key score can be established. Once the first occurrence frequency of the candidate words is determined, the corresponding second key score is determined according to the pre-defined mapping relationship.

[0128] By introducing a first key score and a second key score, and weighting the first key score and the second key score, the target key score of candidate words can be obtained from different dimensions, thereby improving the rationality and accuracy of the target key score. Subsequently, candidate keywords can be determined from the candidate words based on the target key score, which can effectively improve the rationality and accuracy of the candidate keyword determination.

[0129] In one possible implementation, refer to Figure 4 , Figure 4 This is a schematic diagram of an optional process for extracting candidate keywords provided in an embodiment of this application. When constructing a word network based on the adjacency relationship between candidate words in the candidate content and determining the first key score of the candidate words based on the word network, the words in the candidate content can be filtered by part of speech to obtain multiple candidate words; a sliding window of preset width is used to slide the multiple candidate words, and the candidate words located in the sliding window during the sliding process are connected to obtain the word network; the first key score of the candidate words is determined based on the weight of the candidate words in the word network.

[0130] This process involves filtering words in the candidate content by part-of-speech tagging, retaining only words with specific parts of speech as candidate words. This can be done by filtering out words with parts of speech like prepositions, conjunctions, and quantifiers, which are less relevant to the central idea of ​​the content. The remaining words are the candidate words. For example, if the candidate content is "Company X is preparing to hold a groundbreaking ceremony in region Y," then after word segmentation, the resulting words could be "Company X," "preparing," "in," "region Y," "conducting," and "groundbreaking ceremony." After part-of-speech tagging, the candidate words could be "Company X," "region Y," and "groundbreaking ceremony." It is evident that the candidate words in this example are more refined than those in the previous examples, allowing for a more compact and logical word network structure.

[0131] The sliding window is a dynamic window, and its preset width can be determined according to the actual situation, such as 2 or 3, etc., which is not limited in this embodiment. Furthermore, the sliding window can have a fixed width or a variable width. Using the sliding window to process multiple candidate words is actually using the sliding window to traverse multiple candidate words.

[0132] For example, refer to Figure 5 , Figure 5 This is an optional flowchart illustrating a sliding process for multiple candidate words provided in an embodiment of this application. Assuming the candidate content is "Company X is preparing to hold a groundbreaking ceremony in region Y, and Company X will release a brand new product Z on the same day," the candidate words obtained after part-of-speech filtering are "Company X," "region Y," "groundbreaking ceremony," "Company X," "the same day," "release," "brand new," and "product Z." Assuming the width of the sliding window is 2, during the sliding process, the candidate words "Company X" will be connected to the candidate word "region Y," "region Y" will be connected to the candidate word "groundbreaking ceremony," "groundbreaking ceremony" will be connected to the candidate word "Company X," and so on. Therefore, candidate words can be used as nodes, and the relationships between candidate words appearing together in the sliding window can be used as edges to construct a word network.

[0133] For example, refer to Figure 6 , Figure 6 This is a schematic diagram of an optional structure of a word network provided in an embodiment of this application, based on Figure 5 The sliding operation shown in the example ultimately yields a word network with nodes for "Company X", "Region Y", "Opening Ceremony", "Company X", "That Day", "Release", "New", and "Product Z".

[0134] After constructing the word network, a corresponding weight can be preset for each connection. For example, the weight of each connection can be set to 1. Then, based on the number of connections to each node and the weight of each connection, the weight of the candidate word in the word network can be determined. In one possible implementation, for a certain candidate word, the first key score of the candidate word can be obtained by summing the weights of the other candidate words connecting to it.

[0135] Refer again Figure 4 In one possible implementation, when determining the second key score of candidate words based on the first frequency of occurrence, specifically, the number of candidate words can be determined, and a first frequency score can be obtained based on the quotient between the first frequency of occurrence and the number of candidate words; the first content quantity of all candidate content and the second content quantity of candidate content containing candidate words can be determined, and a second frequency score can be obtained based on the quotient between the first content quantity and the second content quantity; and the second key score of candidate words can be obtained based on the first frequency score and the second frequency score.

[0136] The number of candidate words is the total number of words in the candidate content. The first frequency score, obtained by dividing the first occurrence frequency by the number of words, can indicate the importance of a candidate word in a certain candidate content.

[0137] For example, if the candidate content is "Company X is preparing to hold a groundbreaking ceremony in region Y, and all of Company X's senior executives will attend the ceremony", then the first occurrence frequency of the candidate word "Company X" is 2, while the number of words in the candidate content (without part-of-speech filtering) is 13, so the first frequency score is 0.15.

[0138] The second frequency score, obtained by calculating the quotient between the first and second content quantities, indicates the prevalence of candidate words across all candidate content.

[0139] For example, if there are candidate content H1, candidate content H2 and candidate content H3, and a certain candidate word only appears in candidate content H1, then the second frequency score of that candidate word is 3.

[0140] In one possible implementation, the second frequency score is obtained by calculating the logarithm of the quotient between the first and second content quantities. Calculating the base-10 logarithm of this quotient serves a similar purpose to enhancing discriminative power, spreading out crowded values ​​as much as possible and converging outliers, thereby improving the accuracy of the second frequency score.

[0141] Finally, based on the examples of the first and second frequency scores above, the second criticality score is 0.15 * 3 = 0.45.

[0142] By introducing a first frequency score and a second frequency score to calculate a second key score, the frequency of occurrence of a candidate word in a certain candidate content and in all candidate contents can be combined to determine the category discrimination ability of the candidate word itself, thereby making the second key score more accurate and reasonable.

[0143] In a possible implementation, referring to Figure 7 , Figure 7 FIG. is an optional flowchart for extracting candidate entities provided by an embodiment of the present application. When extracting candidate entities in candidate contents, it can be specifically processed at the sentence level and the content level.

[0144] At the sentence level, initial entities in the candidate content can be extracted, and the second occurrence frequency of the initial entities in the reference content can be determined; based on the recognition method corresponding to the second occurrence frequency, ambiguity recognition is performed on the initial entities to obtain the ambiguity recognition result of the initial entities; according to the ambiguity recognition result, candidate entities of the candidate content are determined from the initial entities.

[0145] Among them, the initial entities in the candidate content can be extracted by means of dictionary matching, based on the full name recognition or the short name recognition method. And the short name of a certain entity can be mined by heuristic rules.

[0146] Among them, the reference content is content of a different category from the candidate content. For example, if the candidate content H1 is content related to securities, the candidate content H2 is content related to literature, and the candidate content H3 is content related to science, then the candidate content H2 and the candidate content H3 are the reference contents corresponding to the candidate content H1, and the second occurrence frequency of a certain initial entity in the candidate content H1 is the occurrence frequency of the initial entity in the candidate content H2 and the candidate content H3. Since the reference content and the candidate content are of different categories, if an initial entity in the candidate content appears in the reference content, then the initial entity may be ambiguous. For example, the entity "Murraya paniculata" may refer to a plant or a novel; and for another example, the short name or the full name of the initial entity may also cause ambiguity. Therefore, after the initial entities in the candidate content are extracted, the ambiguity recognition of the initial entities can be performed according to the second occurrence frequency of the initial entities in the reference content to obtain the ambiguity recognition result of the initial entities; determining candidate entities of the candidate content from the initial entities according to the ambiguity recognition result can effectively improve the reliability and accuracy of the candidate entities.

[0147] Among them, the ambiguity recognition result of the initial entity can be used to indicate that the initial entity is an ambiguous entity in this candidate content, or to indicate that the initial entity is a non-ambiguous entity in this candidate content.

[0148] On this basis, further perform ambiguity recognition on the initial entity based on the recognition method corresponding to the second occurrence frequency, that is, different recognition methods are used for different second occurrence frequencies, which is beneficial to improving the refinement degree of ambiguity recognition of the initial entity and enhancing the accuracy of ambiguity recognition of the initial entity.

[0149] In a possible implementation manner, perform ambiguity recognition on the initial entity based on the recognition method corresponding to the second occurrence frequency. Specifically, when the second occurrence frequency is greater than the first frequency threshold, ambiguity recognition is performed on the initial entity according to the context information of the initial entity in the candidate content.

[0150] Specifically, the context information of the initial entity in the candidate content can be used to indicate the semantics expressed by the initial entity in the candidate content. The context information can include other entities or words that appear in the context of the candidate content of the initial entity. When there are other entities of the same category as the initial entity, or words related to the initial entity, or the full name, abbreviation or alias of the initial entity appear in the context information, it can be determined that the initial entity is a non-ambiguous entity in this candidate content.

[0151] For example, assume the initial entity is "Qilixiang". If the full name "Qilixiang Group" of the initial entity appears in the context of the candidate content, it indicates that the initial entity being "Qilixiang" is likely to conform to the context of the candidate content, and it can be determined that the initial entity being "Qilixiang" is a non-ambiguous entity in this candidate content.

[0152] That is, when the second occurrence frequency is greater than the first frequency threshold, it indicates that the ambiguity level of the initial entity is relatively high. Therefore, through the context information of the initial entity in the candidate content, it is possible to effectively and reasonably perform ambiguity recognition on the initial entity.

[0153] In a possible implementation manner, when the second occurrence frequency is greater than the first frequency threshold, even if it is determined that the initial entity is a non-ambiguous entity, when performing relevance matching using the initial entity subsequently, a preset word list can be used to correct the initial entity, and then perform relevance matching using the corrected initial entity. Specifically, the preset word list can store the mapping relationship between the standardized description of an entity and its abbreviation or alias. Therefore, when an initial entity needs to be recognized for ambiguity due to being an abbreviation or alias, even if the initial entity is determined to be a non-ambiguous entity, the standardized description of the initial entity can be matched through the word list subsequently. For example, if the initial entity is "Qilixiang", when performing relevance matching using this initial entity, the initial entity "Qilixiang" can be corrected to "Qilixiang Group" based on the word list.

[0154] When the frequency of the second occurrence is greater than or equal to the second frequency threshold, and the frequency of the second occurrence is less than the first frequency threshold, the initial entity is input into the pre-trained classification model to perform ambiguity recognition on the initial entity.

[0155] In this model, the first frequency threshold is greater than the second frequency threshold. The classification model is used to identify whether the initial entity is ambiguous. After the initial entity is input into the classification model, it can be vectorized to obtain the entity vector. Then, the entity vector is mapped using a preset text encoding to obtain the semantic vector of the initial entity. The semantic vector is input into the ambiguity classifier of the classification model, and the ambiguity classification result of the initial entity is output. This ambiguity classification result can be used to indicate whether the initial entity is ambiguous or unambiguous. When the second occurrence frequency is greater than or equal to the second frequency threshold and less than the first frequency threshold, it indicates that the ambiguity level of the initial entity is not high. Using a pre-trained ambiguity model to identify the ambiguity of the initial entity can improve the efficiency of ambiguity identification.

[0156] In one possible implementation, in order to improve the accuracy of the classification model, in addition to setting an ambiguity classifier, a part-of-speech classifier can also be set in the classification model. The part-of-speech classifier is used to determine whether an initial entity belongs to a real entity, thereby verifying the authenticity of the initial entity input into the classification model and improving the reliability of the classification model.

[0157] In addition, the classification results of the part-of-speech classifier and the ambiguity classifier can be combined and output.

[0158] In one possible implementation, when training the above classification model, sample entities labeled with ambiguous tags can be obtained. The ambiguous tags are used to indicate whether the sample entity is an ambiguous entity or an unambiguous entity. Then, the sample entities are input into the classification model to obtain the ambiguous classification result of the sample entities. Based on the ambiguous classification result of the sample entities and the corresponding ambiguous tags, the first loss value of the classification model is calculated. Then, the parameters of the classification model are adjusted based on the first loss value. The first loss value can be the cross-entropy loss value.

[0159] When a part-of-speech (POS) classifier is included in the classification model, sample entities can also be labeled with POS tags. These tags indicate that the sample entity is a true entity. In this case, sample words labeled with POS tags can also be obtained. These tags indicate that the sample word is not a true entity; for example, the word might be a conjunction or a time term. Next, both the sample entities and sample words are input into the classification model to obtain the POS classification results for the sample entities. Based on the POS classification results and their corresponding POS tags, a second loss value for the classification model is calculated. The parameters of the classification model are then adjusted based on this second loss value. The second loss value can be either cross-entropy loss or triplet loss.

[0160] In addition, the first loss value and the second loss value can be weighted to obtain the target loss value. The parameters of the classification model can be adjusted according to the target loss value to improve the training effect of the classification model.

[0161] Refer again Figure 7 At the content level, entity analysis is first performed. When determining candidate entities for candidate content from the initial entities based on the ambiguity identification results, specifically, ambiguous entities can be identified from the initial entities based on the ambiguity identification results. Initial entities other than ambiguous entities are taken as entities to be determined. Based on the entity position of the entity to be determined in the candidate content, the first entity score of the entity to be determined is determined. The number of paragraphs in the candidate content in which the entity to be determined appears is determined, and the second entity score of the entity to be determined is determined based on the number of paragraphs. The first entity score and the second entity score are weighted to obtain the target entity score of the entity to be determined. Based on the target entity score, candidate entities for candidate content are determined from the entities to be determined.

[0162] After identifying ambiguous entities, the initial entities remaining after removing them are the entities to be determined. The entity position of the entity to be determined within the candidate content can be a paragraph or a title. Different entity positions within the candidate content correspond to different first entity scores. Specifically, the first entity score when the entity to be determined is located in the title of the candidate content can be higher than the first entity score when it is located in a paragraph. When determining the first entity score, a mapping relationship between entity position and first entity score can be preset. After determining the entity position within the candidate content, the corresponding first entity score is determined according to the preset mapping relationship. Similarly, the number of paragraphs containing the entity to be determined in the candidate content can be counted. Different numbers of paragraphs correspond to different second entity scores, with a higher number of paragraphs containing the entity to be determined resulting in a higher second entity score. When determining the second entity score, a mapping relationship between the number of paragraphs containing the entity to be determined can be preset. After determining the number of paragraphs containing the entity to be determined in the candidate content, the corresponding second entity score is determined according to the preset mapping relationship.

[0163] After identifying ambiguous entities, further determining candidate entities from the undetermined entities based on the target entity score makes the final extracted candidate entities more accurate and reliable, improving the accuracy and reliability of subsequent relevance matching. Furthermore, introducing first and second entity scores to determine the target entity score from multiple dimensions makes the target entity score more accurate and reliable, further improving the accuracy and reliability of the final extracted candidate entities.

[0164] Based on this, when determining candidate entities for candidate content from entities to be determined according to the target entity score, entities to be screened can be determined from entities to be determined according to the target entity score, and then candidate entities for candidate content can be screened from entities to be screened according to the preset screening strategy.

[0165] Specifically, the preset filtering strategy is used to identify irrelevant entities from the entities to be filtered, and the remaining entities after removing irrelevant entities are selected as candidate entities for candidate content. The process of filtering candidate entities for candidate content according to the preset filtering strategy can specifically involve determining whether a candidate entity is followed by a preset connecting word or symbol connecting it to another candidate entity. For example, the connecting word could be a word indicating a conclusion, such as "believes" or "feels," and the connecting symbol could be a colon, etc. Taking securities as an example, if the candidate entity is "Securities M," and the relevant text in the candidate content is "Securities M believes: Company N has performance problems," then this candidate content mainly describes the entity "Company N." Therefore, the candidate entity "Securities M" can be removed and not considered as a candidate entity for candidate content.

[0166] In addition, candidate entities for candidate content can be selected from the entities to be screened according to the preset screening strategy, or irrelevant entities can be identified by a pre-trained semantic recognition model, and the remaining entities after removing irrelevant entities from the entities to be screened can be used as candidate entities for candidate content.

[0167] When determining candidate entities for candidate content from entities to be determined based on the target entity score, the candidate entities can be more accurate and reliable by determining the entities to be screened from the entities to be determined based on the target entity score and then filtering the candidate entities for candidate content from the entities to be screened according to the preset filtering strategy.

[0168] It is understandable that the principle of extracting historical keywords and entities from historical content is similar to that of extracting candidate keywords and entities from candidate content, and will not be elaborated here.

[0169] In one possible implementation, when obtaining multiple candidate contents to be pushed, multiple contents to be matched can be obtained; keywords or entities to be matched are extracted from the contents to be matched, and the multiple contents to be matched are classified according to the keywords or entities to be matched to obtain the content index of the multiple contents to be matched under different keywords or entities; historical keywords or historical entities are extracted from historical content, and multiple candidate contents to be pushed are obtained based on the content index corresponding to the historical keywords or historical entities.

[0170] The principle of extracting keywords and entities to be matched from the content to be matched is similar to that of extracting candidate keywords and candidate entities from the candidate content, and will not be repeated here. Each keyword or entity to be matched has a corresponding content index. When there are multiple keywords or entities to be matched, the number of content indexes will also be multiple.

[0171] For example, refer to Figure 8 , Figure 8 This is an optional structural diagram of the content index provided in an embodiment of this application. Assuming the keywords to be matched are "keyword T1" and "keyword T2", and the entities to be matched are "entity S1" and "entity S2", then the content to be matched containing "keyword T1" (a1 to an) forms the content index corresponding to "keyword T1"; the content to be matched containing "keyword T2" (b1 to bn) forms the content index corresponding to "keyword T2"; the content to be matched containing "entity S1" (c1 to cn) forms the content index corresponding to "entity S1"; and the content to be matched containing "entity S2" (d1 to dn) forms the content index corresponding to "entity S2". It can be understood that the content to be matched contained in the content indices corresponding to "keyword T1", "keyword T2", "entity S1", and "entity S2" can overlap.

[0172] In addition, the content index includes content identifiers for a specific piece of content to be matched. These identifiers can be sorted based on the relevance between the content to be matched and the corresponding keywords or entities in the content index, or based on the publication time of the content to be matched. After extracting historical keywords or entities from historical content, the corresponding content index can be determined based on these keywords or entities. Then, multiple content identifiers for content to be matched can be obtained from the corresponding content index, and the corresponding content to be matched can be obtained based on these identifiers as multiple candidate contents to be pushed. If the content identifiers have already been sorted in the aforementioned manner, content identifiers ranked before a preset third ranking threshold can be obtained. This third ranking threshold can be set according to actual conditions, such as 5 or 10, and is not limited in this embodiment.

[0173] It is understood that the specific form of the content index can be determined according to the actual situation. For example, it can be the title of the content to be matched, or other predefined strings. This application does not limit this.

[0174] In addition, whenever new content to be matched is published, the newly published content can be retrieved, and the keywords or entities to be matched in the newly published content can be extracted. The newly published content to be matched can then be categorized based on the keywords or entities to achieve real-time updates to the content index.

[0175] By constructing content indexes for multiple content to be matched under different matching keywords or matching entities, the efficiency of retrieving candidate content from the content to be matched can be effectively improved.

[0176] In one possible implementation, when there are multiple target contents, the target contents are pushed. Specifically, the target contents are structurally analyzed to obtain multiple structural features; an initial structural score is determined based on each structural feature; multiple initial structural scores are weighted to obtain a target structural score; multiple target contents are filtered based on the target structural score, and the filtered target contents are pushed.

[0177] Structural features are used to indicate the content structure information of the target content. These features can be at least two of the following: word count, text-to-image ratio, paragraph structure, and number of keywords. Furthermore, different structural features correspond to different initial structural scores. When determining the initial structural score, a mapping relationship between structural features and initial structural scores can be preset. Once the structural features of the target content are determined, the corresponding initial structural score is determined based on the preset mapping relationship. Different structural features can also correspond to different weights. After determining the initial structural scores corresponding to different structural features, multiple initial structural scores are weighted according to their respective weights to obtain the target structural score.

[0178] For example, for a specific target content, if the initial structure score corresponding to the word count is 0.1, and the weight of this initial structure score is 0.3; the initial structure score corresponding to the text-image ratio is 0.2, and the weight of this initial structure score is 0.3; the initial structure score corresponding to the paragraph structure is 0.1, and the weight of this initial structure score is 0.4; then the target structure score is 0.1*0.3 + 0.2*0.3 + 0.1*0.4 = 0.13. It is understood that the weights corresponding to different structural features can be determined according to the actual situation, and this application does not impose any limitations.

[0179] After obtaining the target structure score, multiple target contents can be sorted according to the target structure score, and target contents ranked before a preset fourth ranking threshold can be pushed. The fourth ranking threshold can be set according to the actual situation, such as 5 or 10, etc., and is not limited in this embodiment. Alternatively, target contents with a target structure score greater than or equal to the preset score threshold can be pushed.

[0180] By filtering multiple target content items based on their target structure scores and then pushing the filtered content, the quality of the pushed target content can be effectively improved. Furthermore, when determining the target structure score, calculating initial structure scores corresponding to multiple different structural features and then weighting these initial structure scores to obtain the final target structure score can effectively improve the accuracy and reliability of the target structure score, thereby optimizing the effect on improving the quality of the pushed target content.

[0181] In one possible implementation, when there are multiple target contents, the target contents are pushed. Specifically, the title similarity and body text similarity between any two target contents can be determined, the title similarity and body text similarity can be weighted to obtain the content similarity, and the multiple target contents can be filtered based on the content similarity, and the filtered target contents can be pushed.

[0182] This can be achieved by vectorizing the titles of the target content to obtain title vectors, and then determining the title similarity between two target content items based on parameters such as cosine similarity or Euclidean distance between their title vectors. Similarly, the target content can be vectorized to obtain text vectors, and then determining the text similarity between two target content items based on parameters such as cosine similarity or Euclidean distance between their text vectors. Title similarity and text similarity each have their own corresponding weights, with the weight of title similarity potentially being less than that of text similarity. It is understood that the weights corresponding to title similarity and text similarity can be determined based on actual circumstances, and this application does not impose any limitations.

[0183] By filtering multiple target content items based on content similarity and then pushing the filtered content, the quality of the pushed target content can be effectively improved. Furthermore, when determining content similarity, weighting the title similarity and body text similarity scores can effectively improve the accuracy and reliability of content similarity measurements, thereby optimizing the effect on improving the quality of the pushed target content.

[0184] In addition, the popularity of each target content can be determined. Popularity indicates how frequently the target content is accessed, and it can be calculated by obtaining the total number of times the target content is accessed within a unit of time. After obtaining multiple target content items, the popularity of each item can be calculated, and the target content can be filtered based on its popularity before being pushed to users.

[0185] Specifically, multiple target contents can be sorted according to popularity, and target contents ranked before a preset fifth ranking threshold can be pushed. The fifth ranking threshold can be set according to the actual situation, such as 5 or 10, etc., and this application embodiment does not limit it.

[0186] It is understandable that the above-mentioned methods of filtering target content based on target structure score, content similarity, and popularity can be combined, selecting at least two indicators from these three factors to filter target content. Taking the combination of target structure score and content similarity as an example, target content can be first filtered based on target structure score, and then further filtered based on content similarity from the results of this two-stage filtering process before being pushed to users. Alternatively, target content can be first filtered based on content similarity, and then further filtered based on target structure score from the results of this two-stage filtering process before being pushed to users. These methods can effectively improve the refinement of target content filtering, thereby further improving the quality of the pushed target content.

[0187] The principle of the content push method provided in this application embodiment is illustrated below with a complete example.

[0188] Reference Figure 9 , Figure 9 This is a schematic diagram of an optional complete process of the content push method provided in this application embodiment. Taking financial news as an example, and historical content and candidate content, multiple candidate contents published within a preset time range (e.g., within 7 days) are obtained. Keyword analysis and entity analysis are performed on the candidate contents. In this example, entity analysis mainly includes related securities analysis, related industry analysis, related event analysis, and keyword analysis. Accordingly, the securities entities, industry entities, event entities, and keywords in the candidate contents can be obtained. Similarly, keyword analysis and entity analysis can be performed on the currently pushed historical content to obtain the securities entities, industry entities, event entities, and keywords in the historical content. Then, based on the securities entities, industry entities, event entities, and keywords in the historical content and candidate content, relevance matching is performed using different matching strategies.

[0189] Specifically, different matching strategies employ different matching benchmark data, which is composed of the aforementioned securities entities, industry entities, event entities, and keywords. In this example, six different matching strategies are combined, and each strategy is assigned a priority based on the matching benchmark data. Different priorities correspond to different relevance weights; the more data types in the matching benchmark data, the higher the relevance weight of the corresponding matching strategy. When the number of data types in the matching benchmark data is the same, the relevance weight of the corresponding matching strategy can be further determined based on the preset priority of the specific data types within the matching benchmark data.

[0190] The aforementioned industry entities are those that appear in historical and candidate content. In this example, we further obtain the industry entities corresponding to the securities entities. For example, if an industry entity appearing in the candidate content is the new energy industry and a securities entity is XX Chemical Co., Ltd., then when combining and matching benchmark data, the chemical industry corresponding to XX Chemical Co., Ltd. can also be used as an industry entity in the combined and matching benchmark data, thereby expanding the matching benchmark data and improving its diversity.

[0191] Therefore, refer to Figure 10 , Figure 10 This is a schematic diagram illustrating an optional division of the matching strategy provided in this application embodiment. In this example, the matching benchmark data corresponding to the first-level matching strategy includes securities entities, industry entities, and keywords; the matching benchmark data corresponding to the second-level matching strategy includes securities entities and keywords; the matching benchmark data corresponding to the third-level matching strategy includes industry entities and keywords; the matching benchmark data corresponding to the fourth-level matching strategy includes industry entities corresponding to securities entities and keywords; the matching benchmark data corresponding to the fifth-level matching strategy includes event entities and keywords; and the matching benchmark data corresponding to the sixth-level matching strategy includes keywords. It is understood that the combination of matching benchmark data and the hierarchical method of the matching strategy in this example can be adjusted according to actual circumstances, and this application embodiment does not impose any limitations.

[0192] Next, using the matching benchmark data corresponding to different matching strategies, the relevance of candidate content and historical content is matched. Taking the first-level matching strategy as an example, the securities entities, industry entities, and keywords of historical content are used as the text set of historical content, and the securities entities, industry entities, and keywords of candidate content are used as the text set of candidate content. The relevance between the text sets of historical content and candidate content is calculated, which is the initial relevance between candidate content and historical content under the first-level matching strategy. In addition, the relevance between securities entities in historical content and securities entities in candidate content, the relevance between industry entities in historical content and industry entities in candidate content, and the relevance between keywords in historical content and keywords in candidate content can also be calculated separately. The average of the relevance between securities entities, the relevance between industry entities, and the relevance between keywords is used as the initial relevance. The processing principles of other levels of matching strategies are similar and will not be elaborated here.

[0193] After determining the initial relevance between historical content and candidate content under various matching strategies, the initial relevance under multiple matching benchmark data is weighted according to the relevance weights under various matching strategies to obtain the target relevance between candidate content and historical content. The target content is then determined from multiple candidate contents based on the target relevance.

[0194] Next, structural analysis is performed on the target content to obtain multiple structural features. An initial structural score for the target content is determined based on each structural feature. Multiple initial structural scores are weighted to obtain the target structural score. Multiple target contents are then screened based on the target structural score to filter out low-quality target content.

[0195] Next, the content similarity between the remaining target content after filtering based on the target structure score is further determined, and deduplication is performed based on the content similarity to obtain the final target content to be pushed.

[0196] Next, after obtaining the final target content to be pushed, the final target content can be sorted from newest to oldest based on publication time or from highest to lowest target relevance, and the top 5 target contents will be pushed.

[0197] In this example, relevance matching can be performed in descending order of relevance weights for different matching strategies. After completing the relevance matching for a certain level of matching strategy, the number of candidate contents with an initial relevance greater than or equal to a preset relevance threshold can be determined. If the number of candidate contents meeting the above conditions is greater than or equal to the preset threshold, relevance matching for subsequent levels of matching strategies can be stopped, thereby improving the efficiency of relevance matching. If the number of candidate contents meeting the above conditions is less than the threshold, relevance matching for subsequent levels of matching strategies continues, ensuring that the target content obtained subsequently has a certain scale, thereby improving the reliability of relevance matching.

[0198] Reference Figure 11 , Figure 11 This is an optional schematic diagram of the content push interface provided in the embodiments of this application. After determining the final target content to be pushed, for example, if 5 target content items are determined, these 5 target content items are displayed on the page of the corresponding securities entity in descending order of target relevance. This can effectively enrich the types of target content, improve the quality and diversity of content push, and improve the efficiency of browsing content.

[0199] It is understood that the above example uses financial news as an example of historical content and candidate content for illustration. In fact, the content push method provided in this application embodiment can also be applied to other types of content. For example, historical content and candidate content can also be sports news. Accordingly, what is extracted from historical content and candidate content are team entities, regional entities, event entities and keywords. The principle of relevance matching is similar to the aforementioned financial news example, and will not be repeated here.

[0200] Accordingly, refer to Figure 12 , Figure 12This is another optional schematic diagram of the content push interface provided in the embodiments of this application. Similarly, after determining the final target content to be pushed, for example, if 5 target content items are determined, these 5 target content items are displayed on the page of the corresponding team entity in descending order of target relevance. This can effectively enrich the types of target content, improve the quality and diversity of content push, and improve the efficiency of browsing content.

[0201] It is understood that although the steps in the above flowcharts are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated in this embodiment, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the above flowcharts may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.

[0202] Reference Figure 13 , Figure 13 This is a schematic diagram of an optional structure of the content push device provided in the embodiments of this application. The content push device 1300 includes:

[0203] The data extraction module 1301 is used to obtain multiple candidate contents to be pushed, extract candidate keywords from the candidate contents, extract candidate entities from the candidate contents, and combine candidate keywords and candidate entities into multiple matching benchmark data.

[0204] The relevance matching module 1302 is used to obtain the historical content that has been pushed, and to perform relevance matching between candidate content and historical content based on various matching benchmark data to obtain the initial relevance between candidate content and historical content under various matching benchmark data.

[0205] The relevance calculation module 1303 is used to weight the initial relevance under multiple matching benchmark data to obtain the target relevance between candidate content and historical content;

[0206] The push module 1304 is used to determine the target content from multiple candidate contents based on the target relevance and push the target content.

[0207] Furthermore, the aforementioned data extraction module 1301 is specifically used for:

[0208] A word network is constructed based on the adjacency relationships between candidate words in the candidate content, and the first key score of the candidate words is determined based on the word network;

[0209] The frequency of the first occurrence of candidate words in candidate content is counted, and the second key score of candidate words is determined based on the frequency of the first occurrence.

[0210] The first and second key scores are weighted to obtain the target key score of the candidate words. Candidate keywords are then determined from the candidate content based on the target key score.

[0211] Furthermore, the aforementioned data extraction module 1301 is specifically used for:

[0212] The words in the candidate content are filtered by part of speech to obtain multiple candidate words;

[0213] Multiple candidate words are processed by sliding a window of preset width. The candidate words located in the sliding window during the sliding process are connected to obtain a word network.

[0214] The first criticality score of a candidate word is determined based on its weight in the word network.

[0215] Furthermore, the aforementioned data extraction module 1301 is specifically used for:

[0216] Determine the number of words in the candidate content, and obtain the first frequency score based on the quotient between the first occurrence frequency and the number of words;

[0217] Determine the first content count of all candidate content and the second content count of candidate content containing candidate words. Obtain the second frequency score based on the quotient between the first content count and the second content count.

[0218] The second key score of the candidate words is obtained based on the first frequency score and the second frequency score.

[0219] Furthermore, the aforementioned data extraction module 1301 is specifically used for:

[0220] Extract the initial entity from the candidate content and determine the second occurrence frequency of the initial entity in the reference content, where the reference content is content that is different from the category of the candidate content;

[0221] Based on the identification method corresponding to the second occurrence frequency, the initial entity is ambiguously identified to obtain the ambiguous identification result of the initial entity;

[0222] Candidate entities for candidate content are determined from the initial entities based on the ambiguity recognition results.

[0223] Furthermore, the aforementioned data extraction module 1301 is specifically used for:

[0224] When the frequency of the second occurrence is greater than the first frequency threshold, the initial entity is ambiguously identified based on the context information of the initial entity in the candidate content.

[0225] When the frequency of the second occurrence is greater than or equal to the second frequency threshold, and the frequency of the second occurrence is less than the first frequency threshold, the initial entity is input into the pre-trained classification model to perform ambiguity recognition on the initial entity;

[0226] Among them, the first frequency threshold is greater than the second frequency threshold.

[0227] Furthermore, the aforementioned data extraction module 1301 is specifically used for:

[0228] Based on the ambiguity identification results, determine the ambiguous entities from the initial entities;

[0229] The initial entities other than ambiguous entities are taken as entities to be determined. The first entity score of the entity to be determined is determined based on the entity position of the entity to be determined in the candidate content.

[0230] Determine the number of paragraphs in the candidate content that contain the entity to be determined, and determine the second entity score of the entity to be determined based on the number of paragraphs.

[0231] The scores of the first entity and the second entity are weighted to obtain the target entity score of the entity to be determined. Based on the target entity score, candidate entities for candidate content are determined from the entities to be determined.

[0232] Furthermore, the aforementioned relevance calculation module 1303 is specifically used for:

[0233] The matching weight corresponding to the matching benchmark data is determined based on the number of data types combined to form the matching benchmark data.

[0234] The initial relevance under various matching benchmark data is weighted according to the corresponding matching weights to obtain the target relevance between candidate content and historical content.

[0235] Furthermore, the aforementioned relevance matching module 1302 is specifically used for:

[0236] Retrieve multiple pieces of content to be matched;

[0237] Extract the keywords or entities to be matched from the content to be matched, and classify multiple contents to be matched according to the keywords or entities to be matched to obtain the content index of multiple contents to be matched under different keywords or entities to be matched;

[0238] Extract historical keywords or entities from historical content, and obtain multiple candidate content to be pushed based on the content index corresponding to the historical keywords or entities.

[0239] Furthermore, the number of target contents is multiple, and the aforementioned push module 1304 is specifically used for:

[0240] Structural analysis of the target content yields multiple structural features.

[0241] The initial structure score of the target content is determined based on each structural feature;

[0242] The target structure score is obtained by weighting the scores of multiple initial structures.

[0243] Multiple target contents are filtered based on the target structure score, and the filtered target contents are then pushed out.

[0244] Furthermore, the number of target contents is multiple, and the aforementioned push module 1304 is specifically used for:

[0245] Determine the title similarity and body text similarity between any two target contents;

[0246] The content similarity is obtained by weighting the title similarity and the body text similarity.

[0247] Multiple target content items are filtered based on content similarity, and the filtered target content is then pushed to users.

[0248] The aforementioned content push device 1300 and content push method are based on the same inventive concept. By acquiring multiple candidate contents to be pushed, extracting candidate keywords and candidate entities from the candidate contents, the matching criteria between candidate contents and historical content can be enriched based on the candidate keywords and candidate entities, improving the comprehensiveness and accuracy of the relevance matching between candidate contents and historical content. Furthermore, by combining candidate keywords and candidate entities into multiple matching benchmark data, the relevance matching between candidate contents and historical content can be performed more comprehensively from multiple perspectives, further improving the comprehensiveness and accuracy of the relevance matching between candidate contents and historical content. Subsequently, the initial relevance under multiple matching benchmark data is weighted to obtain the target relevance between candidate contents and historical content. Based on the target relevance, the target content is determined from multiple candidate contents and pushed. Since the target relevance is obtained by weighting the initial relevance under multiple matching benchmark data, the effect of determining the target content in multiple dimensions can be achieved, thereby effectively enriching the types of target content, improving the quality and diversity of content push, and improving the efficiency of content browsing.

[0249] The electronic device provided in this application embodiment for executing the above-described content push method can be a terminal, as shown below. Figure 14 , Figure 14This is a partial structural block diagram of a terminal provided in an embodiment of this application. The terminal includes: a radio frequency (RF) circuit 1410, a memory 1420, an input unit 1430, a display unit 1440, a sensor 1450, an audio circuit 1460, a wireless fidelity (WiFi) module 1470, a processor 1480, and a power supply 1490, etc. Those skilled in the art will understand that... Figure 14 The terminal structure shown does not constitute a limitation on the terminal and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0250] The RF circuit 1410 can be used to receive and transmit signals during information transmission or calls. In particular, it receives downlink information from the base station and processes it with the processor 1480; in addition, it transmits uplink data to the base station.

[0251] The memory 1420 can be used to store software programs and modules. The processor 1480 executes various functional applications and data processing of the terminal by running the software programs and modules stored in the memory 1420.

[0252] The input unit 1430 can be used to receive input numeric or character information, and to generate key signal inputs related to the settings and function control of the terminal. Specifically, the input unit 1430 may include a touch panel 1431 and other input devices 1432.

[0253] Display unit 1440 can be used to display input or provided information, as well as various menus of the terminal. Display unit 1440 may include display panel 1441.

[0254] Audio circuitry 1460, speaker 1461, and microphone 1462 provide an audio interface.

[0255] In this embodiment, the processor 1480 included in the terminal can execute the content push method of the previous embodiment.

[0256] The electronic device provided in this application embodiment for executing the above content push method can also be a server, see below. Figure 15 , Figure 15The diagram illustrates a partial structural block of a server provided in this application embodiment. The server 1500 can vary significantly due to different configurations or performance characteristics. It may include one or more Central Processing Units (CPUs) 1522 (e.g., one or more processors) and a memory 1532, and one or more storage media 1530 (e.g., one or more mass storage devices) for storing application programs 1542 or data 1544. The memory 1532 and storage media 1530 may be temporary or persistent storage. The program stored in the storage media 1530 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the server 1500. Furthermore, the CPU 1522 may be configured to communicate with the storage media 1530 and execute the series of instruction operations in the storage media 1530 on the server 1500.

[0257] Server 1500 may also include one or more power supplies 1526, one or more wired or wireless network interfaces 1550, one or more input / output interfaces 1558, and / or one or more operating systems 1541, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0258] The processor in server 1500 can be used to execute content push methods.

[0259] This application also provides a computer-readable storage medium for storing program code, which is used to execute the content push methods of the foregoing embodiments.

[0260] This application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the content push method described above.

[0261] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatuses.

[0262] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0263] It should be understood that in the description of the embodiments of this application, "multiple" means two or more, "greater than", "less than", "exceeding" etc. are understood to exclude the number itself, and "above", "below", "within" etc. are understood to include the number itself.

[0264] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0265] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0266] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0267] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0268] It should also be understood that the various implementation methods provided in this application can be combined arbitrarily to achieve different technical effects.

[0269] The above provides a detailed description of the preferred embodiments of this application. However, this application is not limited to the above-described embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.

Claims

1. A content push method, characterized in that, include: Obtain multiple candidate contents to be pushed, extract candidate keywords from the candidate contents, extract initial entities from the candidate contents, and determine the second occurrence frequency of the initial entity in reference content, wherein the reference content is content that is different from the category of the candidate contents; When the second occurrence frequency is greater than the first frequency threshold, the initial entity is ambiguously identified based on the context information of the initial entity in the candidate content; when the second occurrence frequency is greater than or equal to the second frequency threshold and the second occurrence frequency is less than the first frequency threshold, the initial entity is input into a pre-trained classification model to perform ambiguity identification on the initial entity; wherein, the first frequency threshold is greater than the second frequency threshold; Based on the ambiguity identification result, ambiguous entities are identified from the initial entities; the initial entities other than the ambiguous entities are taken as entities to be determined, and a first entity score of the entity to be determined is determined based on the entity position of the entity to be determined in the candidate content; the number of paragraphs in the candidate content in which the entity to be determined appears is determined, and a second entity score of the entity to be determined is determined based on the number of paragraphs; the first entity score and the second entity score are weighted to obtain a target entity score of the entity to be determined, and candidate entities of the candidate content are determined from the entities to be determined based on the target entity score; After deduplication of candidate keywords and candidate entities, the candidate keywords and candidate entities are combined into multiple matching benchmark data. Obtain the historical content that has been pushed, calculate the intersection of the matching benchmark data of the candidate content and the matching benchmark data of the historical content to obtain the intersection of matching data, calculate the union of the matching benchmark data of the candidate content and the matching benchmark data of the historical content to obtain the union of matching data, and obtain the initial relevance between the candidate content and the historical content under various matching benchmark data based on the quotient between the number of elements in the intersection of matching data and the number of elements in the union of matching data. Based on the number of data types that make up the matching benchmark data, the matching weight corresponding to the matching benchmark data is determined. The initial relevance under multiple matching benchmark data is weighted according to the corresponding matching weight to obtain the target relevance between the candidate content and the historical content. When the number of data types that make up the matching benchmark data is the same, the priority of the data types that make up the matching benchmark data is determined, and the matching weight corresponding to the matching benchmark data is determined according to the priority. Based on the target relevance, target content is determined from multiple candidate contents, and the target content is pushed.

2. The content push method according to claim 1, characterized in that, The extraction of candidate keywords from the candidate content includes: A word network is constructed based on the adjacency relationships between candidate words in the candidate content, and a first key score for the candidate words is determined based on the word network. The first occurrence frequency of the candidate words in the candidate content is counted, and the second key score of the candidate words is determined based on the first occurrence frequency; The first key score and the second key score are weighted to obtain the target key score of the candidate words, and candidate keywords are determined from the candidate words based on the target key score.

3. The content push method according to claim 2, characterized in that, The step of constructing a word network based on the adjacency relationships between candidate words in the candidate content, and determining the first key score of the candidate words based on the word network, includes: The words in the candidate content are filtered by part of speech to obtain multiple candidate words; Multiple candidate words are slid through a sliding window of a preset width, and the candidate words located in the sliding window during the sliding process are connected to obtain a word network; The first criticality score of the candidate word is determined based on the weight of the candidate word in the word network.

4. The content push method according to claim 2 or 3, characterized in that, The determination of the second key score of the candidate words based on the first frequency of occurrence includes: The number of candidate words is determined, and a first frequency score is obtained based on the quotient between the first occurrence frequency and the number of words. Determine the first content quantity of all the candidate content and the second content quantity of the candidate content containing the candidate words, and obtain the second frequency score based on the quotient between the first content quantity and the second content quantity; The second key score of the candidate word is obtained based on the first frequency score and the second frequency score.

5. The content push method according to claim 1, characterized in that, The process of obtaining multiple candidate content to be pushed includes: Retrieve multiple pieces of content to be matched; Extract the keywords or entities to be matched from the content to be matched, and classify the multiple contents to be matched according to the keywords or entities to be matched to obtain the content index of the multiple contents to be matched under different keywords or entities to be matched; Extract historical keywords or historical entities from the historical content, and obtain multiple candidate contents to be pushed based on the content index corresponding to the historical keywords or historical entities.

6. The content push method according to claim 1, characterized in that, The number of target content items is multiple, and the process of pushing the target content items includes: Structural analysis is performed on the target content to obtain multiple structural features of the target content; The initial structural score of the target content is determined based on each of the structural features. The target structure score is obtained by weighting the scores of the multiple initial structures. Based on the target structure score, multiple target contents are filtered, and the filtered target contents are pushed out.

7. The content push method according to claim 1, characterized in that, The number of target content items is multiple, and the process of pushing the target content items includes: Determine the title similarity and body text similarity between any two of the target contents; The content similarity is obtained by weighting the title similarity and the body text similarity. Multiple target contents are filtered based on the content similarity, and the filtered target contents are then pushed to the system.

8. A content push device, characterized in that, include: A data extraction module is used to acquire multiple candidate contents to be pushed, extract candidate keywords from the candidate contents, extract initial entities from the candidate contents, and determine the second occurrence frequency of the initial entity in reference content, wherein the reference content is content that is not of the same category as the candidate contents; when the second occurrence frequency is greater than a first frequency threshold, the initial entity is ambiguously identified based on the context information of the initial entity in the candidate contents; when the second occurrence frequency is greater than or equal to the second frequency threshold and less than the first frequency threshold, the initial entity is input into a pre-trained classification model to perform ambiguity identification on the initial entity; wherein the first frequency threshold is greater than the second frequency threshold; root Based on the ambiguity identification result, ambiguous entities are determined from the initial entities; the initial entities other than the ambiguous entities are taken as entities to be determined, and a first entity score of the entity to be determined is determined based on the entity position of the entity to be determined in the candidate content; the number of paragraphs in the candidate content in which the entity to be determined appears is determined, and a second entity score of the entity to be determined is determined based on the number of paragraphs; the first entity score and the second entity score are weighted to obtain a target entity score of the entity to be determined, and candidate entities of the candidate content are determined from the entities to be determined based on the target entity score; after deduplication of candidate keywords and candidate entities, the candidate keywords and candidate entities are combined into multiple matching benchmark data. The relevance matching module is used to obtain the historical content that has been pushed, calculate the intersection of the matching benchmark data of the candidate content and the matching benchmark data of the historical content to obtain the intersection of matching data, calculate the union of the matching benchmark data of the candidate content and the matching benchmark data of the historical content to obtain the union of matching data, and obtain the initial relevance between the candidate content and the historical content under various matching benchmark data based on the quotient between the number of elements in the intersection of matching data and the number of elements in the union of matching data. The relevance calculation module is used to determine the matching weight corresponding to the matching benchmark data based on the number of data types that make up the matching benchmark data, and to weight the initial relevance under multiple matching benchmark data according to the corresponding matching weight to obtain the target relevance between the candidate content and the historical content; when the number of data types of the matching benchmark data is the same, the priority of the data types that make up the matching benchmark data is determined, and the matching weight corresponding to the matching benchmark data is determined according to the priority. The push module is used to determine the target content from multiple candidate contents based on the target relevance and push the target content.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the content push method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the content push method according to any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the content push method according to any one of claims 1 to 7.