Content recommendation method and device, equipment, medium and program product

By determining account ownership based on account content features and target models, the problem of poor user experience caused by operators frequently publishing similar content is solved, and more personalized content recommendations are achieved.

CN120670646APending Publication Date: 2025-09-19TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410316555.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-19
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In the existing technology, in order to increase popularity, the operating entity uses multiple accounts to publish similar content, resulting in the same object being frequently recommended similar content, affecting the user experience.

Method used

By determining the first account corresponding to the target object's current recommended content, recalling a similar second account based on its content characteristics and the content characteristics of other accounts, and using the target model to measure whether the accounts belong to the same operating entity, it is prohibited to recommend the target account's content to the object within a preset time period.

Benefits of technology

It effectively avoids the frequent push of similar content to the same object and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670646A_ABST
    Figure CN120670646A_ABST
Patent Text Reader

Abstract

The invention provides a content recommendation method and device, equipment, a medium and a program product, and can relate to the artificial intelligence technology, and the method comprises the steps: determining a first account corresponding to the current recommendation content of a target object; based on the P account number content features of the first account number and the P account number content features of other account numbers, M second account numbers similar to the first account number are recalled in the other account numbers; for each second account in the M second accounts, inputting the N account content features of the first account and the N account content features of the second account into a target model to obtain scores of the first account and the second account; based on respective scores of the first account and the M second accounts, determining a target account belonging to the same operation main body as the first account in the M second accounts; and prohibiting recommending the content of the target account to the target object in a preset time period. Therefore, similar contents cannot be frequently pushed to the same object, so that the object experience feeling can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence (AI) technology, and in particular to a content recommendation method, apparatus, device, medium, and program product. Background Art

[0002] With the continuous development of internet technology, single- and multi-modal information flow content, such as videos, images, and text, is becoming increasingly popular. Currently, some operators use multiple accounts to publish similar content to increase popularity. This results in the same audience being frequently recommended similar content, which in turn leads to a poor user experience. Summary of the Invention

[0003] The embodiments of the present application provide a content recommendation method, apparatus, device, medium, and program product, so that the same object will not be frequently pushed similar content, thereby improving the object's experience.

[0004] In a first aspect, an embodiment of the present application provides a content recommendation method, including: determining a first account corresponding to the current recommended content of a target object; based on P account content features of the first account and P account content features of other accounts other than the first account among the accounts that have published content, recalling M second accounts similar to the first account from other accounts; wherein P and M are both positive integers; for each second account in the M second accounts, inputting N account content features of the first account and N account content features of the second account into a target model to obtain scores of the first account and the second account; wherein the score is used to measure whether the first account and the second account belong to the same operating entity; N is a positive integer; based on the respective scores of the first account and the M second accounts, determining a target account that belongs to the same operating entity as the first account from the M second accounts; and prohibiting the content of the target account from being recommended to the target object within a preset time period.

[0005] In a second aspect, an embodiment of the present application provides a content recommendation device, comprising: a determination module, a recall module, an input module and a recommendation module; wherein the determination module is used to determine the first account corresponding to the current recommended content of the target object; the recall module is used to recall M second accounts similar to the first account from other accounts based on P account content features of the first account and P account content features of other accounts other than the first account among the accounts that have published content; wherein P and M are both positive integers; the input module is used to input N account content features of the first account and N account content features of the second account into a target model for each second account in the M second accounts to obtain scores of the first account and the second account; wherein the score is used to measure whether the first account and the second account belong to the same operating entity; N is a positive integer; the determination module is also used to determine a target account from the M second accounts that belongs to the same operating entity as the first account based on the respective scores of the first account and the M second accounts; the recommendation module is used to prohibit recommending the content of the target account to the target object within a preset time period.

[0006] In some implementations, the determination module is specifically configured to: determine K third accounts with scores greater than a target preset score among the M second accounts; wherein K is a positive integer; and determine a target account based on the K third accounts.

[0007] In some implementations, the determination module is specifically used to: determine L fourth accounts among K third accounts whose follower counts are greater than a first preset number; where L is a positive integer; and determine a target account based on the L fourth accounts.

[0008] In some possible implementations, the determination module is specifically used to: select the fifth account with the latest content release time from L fourth accounts; divide the first account into the account cluster to which the fifth account belongs; and determine the account in the account cluster to which the fifth account belongs as the target account.

[0009] In some implementations, the determination module is further configured to: for each of the L fourth accounts, determine the publishing time of the latest content published by the fourth account as the content publishing time of the fourth account.

[0010] In some implementations, the determination module is further used to: for each of the L fourth accounts, determine the average publishing time of the fourth account's most recent Q published content as the content publishing time of the fourth account; where Q is an integer greater than 1.

[0011] In some possible implementations, the device also includes: a sampling module, an acquisition module, and an adjustment module, wherein, before the determination module determines K third accounts with scores greater than the target preset scores among the M second accounts based on the scores of the first account and the M second accounts, the sampling module is used to sample the M second accounts to obtain multiple sampled accounts; the determination module is also used to determine the machine judgment result of whether the first account and the multiple sampled accounts belong to the same operating entity based on the scores of the first account and the multiple sampled accounts and the initial preset scores; the acquisition module is used to obtain the manual judgment result of whether the first account and the multiple sampled accounts belong to the same operating entity; the adjustment module is used to adjust the initial preset score based on the machine judgment result and the manual judgment result to obtain the target preset score.

[0012] In some possible implementations, the device also includes: a training module, wherein the acquisition module is further used to: obtain account pairs and labels corresponding to the account pairs for measuring whether the two accounts in the account pairs belong to the same operating entity; obtain N account content features of each account in the account pair; the determination module is further used to form a training sample with the N account content features of each account in the account pair and the labels corresponding to the account pairs; the training module is used to train the target model based on the training samples.

[0013] In some possible implementations, the recall module is specifically used to: if P is greater than or equal to 2, then for each account content feature among the P account content features, based on the account content feature of the first account and the account content features of the other accounts, recall a second account similar to the first account from other accounts.

[0014] In some possible implementations, the device also includes: a push module, wherein the determination module is further used to determine the similarity between the first account cluster and the second account cluster; wherein the first account cluster and the second account cluster are any two account clusters; the push module is used to: if the similarity between the first account cluster and the second account cluster is greater than a preset similarity, push a prompt message to prompt the R&D personnel to determine whether to merge the first account cluster and the second account cluster.

[0015] In some possible implementations, the determination module is specifically used to: select a sixth account in the first account cluster whose number of followers is greater than a second preset number, and select a seventh account in the second account cluster whose number of followers is greater than a third preset number; and determine the similarity between the first account cluster and the second account cluster based on the account content characteristics of the sixth account and the account content characteristics of the seventh account.

[0016] In some implementations, the determination module is specifically configured to: if the first account cluster includes multiple first candidate accounts whose follower counts are greater than a second preset number, select the first candidate account with the latest content publishing time from the multiple first candidate accounts as the sixth account.

[0017] In some implementations, the determination module is specifically configured to: if the second account cluster includes multiple second candidate accounts whose follower counts are greater than a third preset number, select the second candidate account with the latest content publishing time from the multiple second candidate accounts as the seventh account.

[0018] In some possible implementations, the determination module is specifically used to: convert the account content features of the sixth account and the account content features of the seventh account into the embedding vector of the sixth account and the embedding vector of the seventh account, respectively; calculate the similarity between the embedding vector of the sixth account and the embedding vector of the seventh account to obtain the similarity of any two account clusters.

[0019] In some implementations, the P account content features include at least one of the following: account name, content extraction result.

[0020] In some implementations, the N account content features include at least one of the following: account name, account signature, content title, content frame extraction result, entity recognition result, OCR result, ASR result.

[0021] In a third aspect, an electronic device is provided, comprising: a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method in the first aspect or its various implementations.

[0022] In a fourth aspect, a computer-readable storage medium is provided for storing a computer program, wherein the computer program enables a computer to execute the method according to the first aspect or its various implementations.

[0023] In a fifth aspect, a computer program product is provided, comprising computer program instructions, which enable a computer to execute the method in the first aspect or its various implementations.

[0024] In a sixth aspect, a computer program is provided, which enables a computer to execute the method in the first aspect or its various implementations.

[0025] Through the technical solution provided by this application, since the account content characteristics can represent the account content, the target account belonging to the same operating entity as the first account can be determined more accurately based on the account content characteristics of the first account and the account content characteristics of other accounts, so that the target object will not be frequently pushed similar content, thereby improving the object experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0027] Figure 1 Schematic diagram of videos released by the same operating entity;

[0028] Figure 2 A schematic diagram of a system architecture involved in an embodiment of the present application;

[0029] Figure 3 A flowchart of a content recommendation method provided in an embodiment of the present application;

[0030] Figure 4 A flowchart of a model training method provided in an embodiment of the present application;

[0031] Figure 5 A schematic diagram of account content features corresponding to an account pair provided in an embodiment of the present application;

[0032] Figure 6 A schematic diagram of a content recommendation device 600 provided in an embodiment of the present application;

[0033] Figure 7 It is a schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0035] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0036] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0037] The embodiments of the present application may involve AI technology, but are not limited thereto.

[0038] AI refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, artificial intelligence is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. Artificial intelligence is the study of the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0039] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0040] Computer vision (CV) is the science of making machines "see." Specifically, it refers to machine vision, which uses cameras and computers to replace the human eye in identifying, detecting, and measuring objects. Further image processing is performed to transform the images into images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems that can extract information from images or multidimensional data. Large model technology has brought significant changes to the development of computer vision. Pre-trained models in the field of vision, such as the Swin Transformer, ViT, V-MOE, and MAE, can be fine-tuned to quickly and widely adapt to specific downstream tasks. Computer vision technologies generally include image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / action recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and common biometric recognition technologies such as facial recognition and fingerprint recognition. In an embodiment of the present application, computer vision technology can be used to extract account content features, such as OCR results of videos.

[0041] The key technologies of speech technology include automatic speech recognition (ASR), text-to-speech (TTS) and voiceprint recognition technology. Enabling computers to listen, see, speak and feel is the future development direction of human-computer interaction, among which speech has become one of the most promising ways of human-computer interaction in the future. Large model technology has brought changes to the development of speech technology. Pre-trained models such as WavLM and UniSpeech that use the Transformer architecture have strong generalization and versatility, and can excellently complete speech processing tasks in various directions. In the embodiment of the present application, ASR technology can be used to extract account content features, such as the ASR results of videos.

[0042] Machine Learning (ML) is a multi-disciplinary interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithmic complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning generally include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning. The pre-trained model is the latest development in deep learning and integrates the above technologies. In an embodiment of the present application, machine learning can be used to calculate the scores of two accounts. For example, the account content features of the two accounts can be input into the target model to obtain the scores of the two accounts. The scores are used to measure whether the two accounts belong to the same operating entity. Among them, the target model can be a machine learning model.

[0043] The following is an explanation of the relevant knowledge involved in this application:

[0044] Matrix accounts refer to multiple accounts opened or linked by a single operating entity (either a company or an individual), which direct traffic between accounts to maximize marketing effectiveness as a group. These accounts may include various web-based self-media platforms and short video platforms.

[0045] 2. Multimodal model refers to a model that can process multiple types of data simultaneously, such as text, images, audio and video, or a model whose input data includes multiple types of data.

[0046] 3. Online and offline consistency refers to the situation where, when the same account content features are input into the same model in both offline and online scenarios, the prediction results in both scenarios are consistent or the difference is very small.

[0047] Fourth, an embedding vector is a low-dimensional, dense vector that can be used to represent an object. For example, account content features such as frame extraction results, account name, account signature, etc. can all be represented by corresponding embedding vectors.

[0048] 5. Cluster is a core concept formed in cluster analysis. Clustering is a data analysis method used to group similar objects in a data set into different classes or clusters.

[0049] 6. Entity recognition results: The results obtained by identifying entities in account content using entity recognition methods, such as names of people, places, and organizations.

[0050] 7. OCR result: the result obtained by recognizing text in account content using the OCR method. For example, text in a video frame recognized using the OCR method is an OCR result.

[0051] 8. ASR results: the results obtained by recognizing the voice in the account content using the ASR method.

[0052] 9. Open Neural Network Exchange (ONNX) is a standard format for representing deep learning models, which enables models to be transferred between different deep learning frameworks.

[0053] The following describes the technical problems, inventive concepts, and system architecture to be solved by the embodiments of the present application:

[0054] As mentioned above, some operating entities currently use multiple accounts to publish similar content in order to increase popularity, resulting in the same object being frequently recommended similar content, which in turn leads to a poor user experience for the object.

[0055] For example, Figure 1 Schematic diagram of videos released by the same operating entity, such as Figure 1 The contents of the four video frames shown are similar, and it is highly likely that the accounts publishing these videos belong to the same operating entity.

[0056] In order to solve the above technical problems, this application proposes to determine the target account of the same operating entity as the account corresponding to the current recommended content based on the account content characteristics of other accounts and the account corresponding to the current recommended content, and not push the content of the target account to the object corresponding to the current recommended content within a preset time period, so that the object will not be frequently pushed similar content and affect the experience.

[0057] Figure 2 This is a schematic diagram of a system architecture involved in an embodiment of the present application, including a terminal device 210 and a server 220.

[0058] The terminal device 210 may be installed with a content recommendation application (Application, APP), for example, the content recommendation APP may be a short video APP.

[0059] Alternatively, other APPs may be installed on the terminal device 210, which may provide an interface for entering the content recommendation module. For example, an instant messaging client may be installed on the terminal device 210, which may provide an interface for entering the video account.

[0060] It should be understood that the instant messaging client can be any communication tool based on Internet technology that allows users to communicate with each other in real time through text, voice, video, etc., such as WeChat, QQ, Enterprise WeChat, etc., but not limited to these.

[0061] In some implementations, the terminal device 210 can be a desktop computer, a laptop computer, a tablet computer, a smart phone, a tablet computer, a smart watch, virtual reality (VR), augmented reality (AR), etc., but is not limited to this.

[0062] The server 220 may be a background server corresponding to a content recommendation APP or a content recommendation module.

[0063] In some possible implementations, server 220 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0064] The terminal device 210 and the server 220 may be connected directly or indirectly via wired or wireless communication, which is not limited in this application.

[0065] It should be noted that Figure 2 This is only a schematic diagram of a system architecture provided by the embodiment of the present application. The system architecture involved in the embodiment of the present application is not limited to Figure 2 The system architecture shown, for example, the number of terminal devices 210 is not limited to Figure 2 One is shown, but there may be more than one.

[0066] The following is a detailed description of the embodiments of the present application:

[0067] Figure 3 This is a flowchart of a content recommendation method provided in an embodiment of the present application. The method can be executed by a server, but is not limited thereto. The server can be Figure 2 The server 220 in Figure 3 As shown, the method may include:

[0068] S310: Determine the first account corresponding to the current recommended content of the target object;

[0069] In some implementations, the currently recommended content may be a video, an article, or the like. For example, in a short video app or a video account provided by an instant messaging client, the currently recommended content may be a video. On a web page, the currently recommended content may be an article. This embodiment of the application does not limit the form of the currently recommended content.

[0070] In some possible implementations, if the currently recommended content is a video, the video may be an e-commerce live video or other non-live video, etc., wherein the embodiment of the present application does not limit the playback method of the video.

[0071] In addition, the embodiments of the present application do not limit the specific content of the current recommended content.

[0072] In some implementations, the first account refers to the account used to post the currently recommended content. For example, in a short video app or instant messaging client, the first account may be the short video account used to post the currently recommended content. On a webpage, the first account may be the account of the author of the currently recommended article.

[0073] S320: Based on the P account content features of the first account and the P account content features of accounts other than the first account that have published content, recall M second accounts similar to the first account from other accounts; where P and M are both positive integers;

[0074] In some feasible sendings, the other accounts here may be accounts other than the first account among all accounts that have published content, or may be accounts other than the first account among accounts that have published content within a preset time period, wherein the starting time of the preset time period may be midnight of the day where the currently recommended content is located, and the end time of the preset time period may be the recommendation time when the server recommends the currently recommended content to the target object, or the end time of the preset time period may be the recommendation time when the server recommends the currently recommended content to the target object, and the starting time of the preset time period is 1 hour away from the end time.

[0075] It should be understood that the embodiments of the present application do not limit the preset time period.

[0076] It should be understood that for each of the first account and other accounts, the account information may include, but is not limited to: the account name, the account signature, and one or more recently published content items. For example, for a short video account, the account information may include: the account name, the account signature, and the 10 most recently published videos.

[0077] It should be understood that, for each of the first account and other accounts, the account content feature refers to the feature of the account content of the account.

[0078] In some implementations, the P account content features include at least one of the following: account name, account signature, content title, content frame extraction result, entity recognition result, OCR result, ASR result.

[0079] For each of the first account and the other accounts, the corresponding content title is derived from one or more contents recently published by the account.

[0080] For each of the first account and the other accounts, the corresponding content frame extraction result can be one or more video frames formed by extracting frames from one or more videos recently published by the account, or facial features or other features obtained by performing image recognition on the one or more video frames. The embodiments of the present application do not limit the specific frame extraction method.

[0081] Among them, for each of the first account and the other accounts, the corresponding entity recognition result can be an entity recognition result obtained by performing entity recognition on at least one of the account name, account signature, and content title.

[0082] For each of the first account and the other accounts, the corresponding OCR result may be a result obtained by performing character recognition on the content frame extraction result.

[0083] Among them, for each of the first account and other accounts, the corresponding ASR result can be the result obtained by ASR recognition of one or more videos or voices, etc., which are most recently published through the account.

[0084] It should be understood that, for the first account and any other account, the account content features of the two are corresponding.

[0085] For example, the P account content features of the first account include: the account name of the first account, and correspondingly, the P account content features of other accounts also include: the account names of the other accounts.

[0086] For example, the P account content features of the first account include: the content frame extraction result of the first account, and correspondingly, the P account content features of other accounts also include: the content frame extraction result of the other accounts.

[0087] For example, the P account content features of the first account include: the account name and content extraction results of the first account, and correspondingly, the P account content features of other accounts also include: the account name and content extraction results of the other accounts.

[0088] In some implementations, the server may calculate the similarity between the first account and the other accounts based on the P account content features of the first account and the P account content features of each other account, and use the other accounts with a similarity greater than a preset similarity as the second account.

[0089] In some implementations, if P is greater than or equal to 2, the server may, for each of the P account content features, recall a second account similar to the first account from other accounts based on the account content feature of the first account and the account content features of the other accounts.

[0090] For example, assuming that P account content features include: account name and content extraction results, then the server can recall the second account similar to the first account from other accounts based on the account name of the first account and the account names of the other accounts, and the server can recall the second account similar to the first account from other accounts based on the content extraction results of the first account and the content extraction results of the other accounts, or in other words, the second accounts obtained based on these two account content features are all similar accounts of the first account, or in other words, the union of the sets of second accounts obtained based on the two account content features includes M second accounts.

[0091] In some possible implementations, for each account content feature among P account content features, when the server recalls a second account similar to the first account from other accounts based on the account content feature of the first account and the account content features of other accounts, the server can first convert the account content feature of the first account into a corresponding embedding vector, and replace the account content feature of the other accounts with the corresponding embedding vector. Furthermore, the similarity of the two embedding vectors can be calculated, and the similarity can be used as the similarity between the first account and the other accounts.

[0092] In this embodiment of the present application, the similarity can be cosine similarity, Euclidean distance, Manhattan distance, Pearson correlation coefficient or Jaccard similarity coefficient, etc., and this embodiment of the present application does not limit this.

[0093] It should be understood that S320 can be understood as a rough recall step of similar accounts of the first account, and the following S330 can be understood as a detailed recall step of similar accounts of the first account. The rough recall can narrow the recall scope and improve the recall efficiency of similar accounts.

[0094] S330: For each of the M second accounts, input the N account content features of the first account and the N account content features of the second account into the target model to obtain scores for the first account and the second account; wherein the scores are used to measure whether the first account and the second account belong to the same operating entity; N is a positive integer;

[0095] In some possible implementations, the operating entity may be an individual or an enterprise, etc., and the embodiments of the present application do not impose any restrictions on this.

[0096] It should be understood that, according to the definition of the matrix number, in the embodiment of the present application, the operating entity can also be replaced by the matrix number.

[0097] In some implementations, before using the target model, the server may first deploy it online, then verify the online and offline consistency, and then use the target model after the verification is successful.

[0098] In some implementations, the N account content features include at least one of the following: account name, account signature, content title, content frame extraction result, entity recognition result, OCR result, ASR result.

[0099] It should be understood that the explanation of the content features of N accounts can be found above, and the embodiments of this application will not be repeated here.

[0100] It should be understood that the larger N is, the higher the prediction accuracy of the target model. In particular, when the N account content features are multimodal content features, the prediction accuracy of the target model is higher.

[0101] It should be noted that for the same account, its P account content features and N account content features can be exactly the same, or P equals N. In this case, when the server performs a coarse recall based on the P account content features, it recalls the second account based on each of these account content features. When the server performs a fine recall based on the P account content features, it can input all of these account content features into the target model.

[0102] Alternatively, for the same account, its P account content features and N account content features may not be completely identical, or in other words, P is less than N. For example, for each other account, the P account content features are the account name, and the N account content features include: account name, account signature, content title, content frame extraction results, entity recognition results, OCR results, and ASR results. Based on this, the server can roughly recall M second accounts based on the account name of the first account and the account name of the other account. Furthermore, for each second account, the server inputs the N account content features of the first account and the N account content features of the second account into the target model to obtain the scores of the first account and the second account.

[0103] In some implementations, the target model may be a machine learning model, which is not limited in the embodiments of the present application.

[0104] It should be noted that when the input data of the target model includes multiple types of data, the target model is a multimodal model.

[0105] S340: Based on the scores of the first account and the M second accounts, determine a target account from the M second accounts that belongs to the same operating entity as the first account;

[0106] It should be understood that the target account refers to an account that is predicted to belong to the same operating entity as the first account. The number of target accounts can be one or more.

[0107] In the embodiment of the present application, the target account may be determined by any of the following feasible methods, but is not limited thereto:

[0108] In some implementations, S340 may include:

[0109] S340-1A: Determine K third accounts from the M second accounts whose scores are greater than a target preset score; where K is a positive integer;

[0110] It should be understood that for each second account, the greater the score between the first account and the second account, the more likely it is that the second account and the first account belong to the same operating entity. The smaller the score between the first account and the second account, the less likely it is that the second account and the first account belong to the same operating entity. Based on this, the server first determines, from the M second accounts, K third accounts with scores greater than the target preset score.

[0111] It should be understood that the third account refers to an account among the M second accounts whose corresponding score is greater than the target preset score.

[0112] It should be understood that the target preset score determines the recall of the third account, and thus determines the recall of the target account. Based on this, it determines the accuracy of content recommendation. In order to improve the accuracy of content recommendation, the target preset score needs to be as appropriate as possible. Based on this, the embodiment of the present application proposes the following implementation method:

[0113] In some possible implementations, before determining K third accounts with scores greater than a target preset score among the M second accounts based on the scores of the first account and the M second accounts, the server may sample the M second accounts to obtain multiple sampled accounts; based on the scores of the first account and the multiple sampled accounts and the initial preset scores, determine the machine judgment result of whether the first account and the multiple sampled accounts belong to the same operating entity; obtain the manual judgment result of whether the first account and the multiple sampled accounts belong to the same operating entity; and adjust the initial preset score based on the machine judgment result and the manual judgment result to obtain the target preset score.

[0114] For example, assume M = 1000, the number of sampled accounts is set to 100, and the initial preset score is 6. Suppose the server calculates the scores of each of the 100 sampled accounts relative to the first account. Suppose it is determined that 10 of the 100 sampled accounts have a score greater than 6 with the first account, and 90 of the sampled accounts have a score less than 6 with the first account. Based on this, the machine judgment result is that 10 of the 100 sampled accounts belong to the same operating entity as the first account, and 90 of the sampled accounts do not belong to the same operating entity as the first account. However, using a manual judgment method, 20 of the 100 sampled accounts belong to the same operating entity as the first account, and 80 of the sampled accounts do not belong to the same operating entity as the first account. Based on this, the server can reduce the initial preset score, for example, from 6 to 5.

[0115] For another example, assume M = 1000, the number of sampled accounts is set to 100, and the initial preset score is 6. Suppose the server calculates the scores of each of the 100 sampled accounts relative to the first account. Suppose it is determined that 20 of the 100 sampled accounts have a score greater than 6 with the first account, and 80 of the sampled accounts have a score less than 6 with the first account. Based on this, the machine judgment result is that 20 of the 100 sampled accounts belong to the same operating entity as the first account, and 80 of the sampled accounts do not belong to the same operating entity as the first account. However, using a manual judgment method, 10 of the 100 sampled accounts belong to the same operating entity as the first account, and 20 of the sampled accounts do not belong to the same operating entity as the first account. Based on this, the server can increase the initial preset score, for example, from 6 to 7.

[0116] It should be noted that in order to further improve the accuracy of the target preset score and thus improve the accuracy of content recommendation, the server can sample M second accounts multiple times to obtain multiple groups of sampled accounts. For the first group of sampled accounts, the server can determine the machine judgment result of whether the first account and the multiple sampled accounts in the group belong to the same operating entity based on the respective scores of the first account and the multiple sampled accounts in the group and the initial preset scores; obtain the manual judgment result of whether the first account and the multiple sampled accounts in the group belong to the same operating entity; adjust the initial preset score based on the machine judgment result and the manual judgment result. By analogy, for the second group of sampled accounts, the preset score obtained in the previous round is adjusted in a similar manner to the first group of sampled accounts, until the preset score obtained in the previous round is adjusted through the last group of sampled accounts, and then the process ends.

[0117] In some implementations, the preset score is adjusted based on the machine-based and manual-based results, including: if the difference between the machine-based and manual-based results is less than a preset error, then the server may not adjust the preset score; if the difference between the machine-based and manual-based results is greater than or equal to the preset error, then the server may adjust the preset score. This approach is applicable to both sampling and grouping scenarios and to scenarios where no sampling and grouping is performed.

[0118] In some possible implementations, the server can determine a machine determination result of whether the first account and the M second accounts belong to the same operating entity based on the respective scores of the first account and the M second accounts and the initial preset score; obtain a manual determination result of whether the first account and the M second accounts belong to the same operating entity; and adjust the initial preset score based on the machine determination result and the manual determination result to obtain a target preset score.

[0119] It should be noted that, in this implementable manner, the server does not need to sample the M second accounts.

[0120] S340-2A: Determine a target account based on the K third accounts.

[0121] In the embodiment of the present application, S340-2A can be implemented in any of the following ways, but is not limited thereto:

[0122] In some implementations, S340-2A may include:

[0123] S340-2A-1a: Set the K third accounts as target accounts.

[0124] For example, assuming that the server roughly recalls 1,000 accounts similar to the first account, and the scores of the first account and 100 of the 1,000 accounts are both greater than the target preset score of 6, based on this, the server can use these 100 accounts as target accounts.

[0125] Through this feasible method, the server can recall as many accounts as possible that are suspected to belong to the same operating entity as the first account, and ensure that similar content is not continuously pushed to the same object to ensure the object's experience.

[0126] In some other implementations, S340-2A may include:

[0127] S340-2A-1b: Determine L fourth accounts from the K third accounts whose follower counts are greater than a first preset number; where L is a positive integer;

[0128] In some possible implementations, the number of followers of an account may also be referred to as the number of followers or fans of the account, etc., which is not limited in this embodiment of the present application.

[0129] It should be understood that the fourth account is also called a big account. Correspondingly, for accounts whose number of followers is less than or equal to the first preset number, such accounts can be called small accounts.

[0130] S340-2A-2b: Determine a target account based on the L fourth accounts.

[0131] In the embodiment of the present application, S340-2A-2b can be implemented in any of the following ways, but is not limited thereto:

[0132] In some implementations, S340-2A-2b may include:

[0133] S340-2A-2b-1a: Set L fourth accounts as target accounts.

[0134] For example, assuming that the server has retrieved 100 accounts similar to the first account, and there are two accounts among these 100 accounts whose follower counts are greater than the first preset number of 5,000, based on this, the server can use these two accounts as target accounts.

[0135] Through this feasible method, the server can recall large accounts suspected to belong to the same operating entity as the first account as much as possible, and do not process small accounts suspected to belong to the same operating entity as the first account. This can reduce server power consumption and improve the content recommendation efficiency of the server while ensuring that similar content is not continuously pushed to the same object as much as possible.

[0136] In some other implementations, S340-2A-2b may include:

[0137] S340-2A-2b-1b: Select the fifth account with the latest content publishing time from the L fourth accounts;

[0138] In some implementations, for each of the L fourth accounts, the server may determine the publishing time of the most recently published content by the fourth account as the content publishing time of the fourth account. In other words, for each of the L fourth accounts, the content publishing time of the fourth account refers to the publishing time of the most recently published content by the fourth account.

[0139] For example, suppose an account has published a total of 5 videos. The publishing time of these 5 videos are: 2024-01-01-11:00, 2024-01-03-13:00, 2024-01-10-21:00, 2024-01-12-11:00, 2024-01-15-18:10, then it can be considered that the content of this account was published at 2024-01-15-18:10.

[0140] In some implementations, for each of the L fourth accounts, the average publishing time of the fourth account's most recent Q content releases is determined as the content release time of the fourth account; where Q is an integer greater than 1. In other words, for each of the L fourth accounts, the content release time of the fourth account refers to the average publishing time of the fourth account's most recent Q content releases.

[0141] For example, assuming that an account has published a total of 6 videos, and the publication time of these 6 videos are: 2024-01-01-11:00, 2024-01-01-12:00, 2024-01-01-13:00, 2024-01-01-14:00, 2024-01-01-15:00, 2024-01-01-16:00, then it can be considered that the content release time of this account is approximately equal to 2024-01-01-15:00.

[0142] It should be understood that the reason for selecting the account with the latest content publishing time is that some fourth accounts may have previously published similar content to the first account, but are not currently publishing similar content. If such accounts are selected as target accounts, it may lead to incorrect judgment. In other words, by selecting the fifth account with the latest content publishing time among the L fourth accounts, the recall accuracy of the target account can be improved, and thus the accuracy of content recommendations can be improved, thereby enhancing the user experience.

[0143] S340-2A-2b-2b: assign the first account to the account cluster to which the fifth account belongs;

[0144] It should be understood that the account cluster shown by the current fifth account is divided into two situations:

[0145] Case 1: the account cluster only includes the fifth account.

[0146] In case 2, the account cluster includes other accounts in addition to the fifth account.

[0147] The following describes the second scenario:

[0148] It should be understood that before executing S310, the content recommendation method provided in the embodiment of the present application is also used for other accounts, and then these other accounts may also be divided into the account cluster to which the fifth account belongs.

[0149] S340-2A-2b-3b: Determine an account in the account cluster to which the fifth account belongs as a target account.

[0150] For example, assuming that the first account is recorded as F, its corresponding fifth account is recorded as f_max, and assuming that the account cluster to which the fifth account belongs is {a, b, c, F, f_max}, then accounts a, b, c, F, f_max can all be determined as target accounts.

[0151] In some implementations, the account cluster to which the fifth account belongs may further include: the content publishing time of the fifth account.

[0152] For example, the account cluster to which the fifth account belongs is {a, b, c, F, f_max, 2024-01-01-11:00}, and the time 2024-01-01-11:00 is the content publishing time of the fifth account f_max.

[0153] It should be noted that, as described above, if the other accounts in S310 are accounts other than the first account among the accounts that have published content within the preset time period, then the account cluster to which the fifth account belongs is also the account cluster generated within the preset time period. For example, the account cluster can be at the hourly level.

[0154] It should be understood that the account in the account cluster to which the fifth account belongs is determined as the target account because, according to the content recommendation method provided in the embodiment of this application, all accounts in the account cluster are similar to the fifth account (i.e., the large account with the latest content release time). This achievable method can not only improve the accuracy of content recommendation, but also recall as many accounts as possible that are suspected to belong to the same operating entity as the first account, ensuring that similar content is not continuously pushed to the same person, thereby ensuring the person's experience.

[0155] In some other possible implementations, S340 may include:

[0156] S340-1B: Determine S eighth accounts among the M second accounts whose scores multiplied by a preset coefficient are greater than a target preset score; where S is a positive integer;

[0157] S340-2B: Determine the target account based on the S eighth accounts.

[0158] It should be understood that the eighth account refers to an account among the M second accounts whose corresponding score multiplied by a preset coefficient is greater than the target preset score.

[0159] It should be understood that the difference between this implementation of S340 and the previous implementation is that in this implementation, the server can multiply the score of the second account by a preset coefficient and compare the resulting product with the target preset score, while in the previous implementation, the server directly compares the score of the second account with the target preset score. Therefore, all implementations introduced in the previous implementation also apply to this implementation, and will not be further described in this embodiment of the present application.

[0160] S350: Prohibiting the recommendation of the target account's content to the target object within a preset time period.

[0161] In some possible implementations, the starting time of the preset time period may be the time when the server recommends the current recommended content to the target object, or the time when the target account is determined, etc. The embodiment of the present application does not limit the starting time of the preset time period.

[0162] In some implementations, the duration of the preset time period may be a preset duration, such as 30 minutes, 60 minutes, 24 hours, etc., and the embodiments of the present application do not limit this.

[0163] The present application provides a content recommendation method, comprising: determining a first account corresponding to currently recommended content for a target object; recalling M second accounts similar to the first account from other accounts based on P account content features of the first account and P account content features of accounts other than the first account that have published content; inputting N account content features of the first account and N account content features of the second account into a target model for each of the M second accounts to obtain scores for the first account and the second account; wherein the scores are used to measure whether the first account and the second account belong to the same operating entity; based on the scores of the first account and the M second accounts, determining a target account from the M second accounts that belongs to the same operating entity as the first account; and prohibiting the recommendation of content from the target account to the target object within a preset time period. Because account content features can represent account content, target accounts belonging to the same operating entity as the first account can be more accurately determined based on the account content features of the first account and the account content features of the other accounts, thereby preventing the target object from being frequently pushed similar content, thereby improving the user experience.

[0164] It should be understood that since each account cluster includes a large account, the merger between account clusters includes the merger between large accounts, and the merger between large accounts should be stricter to prevent large accounts that do not actually belong to the same operating entity from being mistakenly identified as belonging to the same operating entity. For example, although a TV station’s video account and a certain operator’s video account have recently been broadcasting the same hot news, the TV station’s video account and a certain operator’s video account cannot be identified as accounts belonging to the same operating entity. Based on this, the embodiment of the present application proposes the following implementable methods:

[0165] In some possible implementations, the server may determine the similarity between the first account cluster and the second account cluster; if the similarity between the first account cluster and the second account cluster is greater than a preset similarity, the server pushes a prompt message to prompt the R&D personnel to determine whether to merge the first account cluster and the second account cluster.

[0166] In some implementations, if the similarity between the first account cluster and the second account cluster is less than or equal to a preset similarity, the server does not perform any processing.

[0167] The first account cluster and the second account cluster are any two account clusters.

[0168] In this embodiment, the similarity between the first account cluster and the second account cluster may be determined by any of the following achievable methods, but is not limited thereto:

[0169] In a first implementation method, the server may select a sixth account from the first account cluster whose number of followers is greater than a second preset number, and select a seventh account from the second account cluster whose number of followers is greater than a third preset number; and determine the similarity between the first account cluster and the second account cluster based on the account content characteristics of the sixth account and the account content characteristics of the seventh account.

[0170] In some possible implementations, the second preset number, the third preset number and the first preset number may be completely the same or may not be completely the same, and the embodiments of the present application do not limit this.

[0171] It should be understood that the sixth account can be understood as the large account in the first account cluster.

[0172] It should be understood that the first account cluster may include one or more accounts with a follow count greater than the second preset number. In this embodiment of the present application, such accounts may be referred to as first candidate accounts. Based on this, if the first account cluster includes one first candidate account, the server may use the first candidate account as the sixth account. If the first account cluster includes multiple first candidate accounts, the server may determine the sixth account using any of the following implementable methods, but is not limited to these:

[0173] In some implementations, the server selects the first candidate account with the latest content publishing time from among multiple first candidate accounts as the sixth account.

[0174] It should be understood that the explanation of the content release time can be found above, and the embodiments of the present application will not be repeated here.

[0175] For example, assuming there are three first candidate accounts 001, 002, and 003, whose content publishing times are: 2024-01-01-11:00, 2024-01-01-12:00, 2024-01-01-13:00, then the server can use 003 as the sixth account.

[0176] In some other implementations, the server randomly selects a first candidate account from multiple first candidate accounts as the sixth account.

[0177] It should be understood that the seventh account can be understood as the large account in the second account cluster.

[0178] It should be understood that the second account cluster may include one or more accounts with a follow count greater than a third preset number. In this embodiment of the present application, such accounts may be referred to as second candidate accounts. Based on this, if the second account cluster includes one second candidate account, the server may use this second candidate account as the seventh account. If the second account cluster includes multiple second candidate accounts, the server may determine the seventh account using any of the following implementable methods, but is not limited to these:

[0179] In some implementations, the server selects the second candidate account with the latest content publishing time from among the multiple second candidate accounts as the seventh account.

[0180] It should be understood that the explanation of the content release time can be found above, and the embodiments of the present application will not be repeated here.

[0181] In some other implementations, the server randomly selects a second candidate account from multiple second candidate accounts as the seventh account.

[0182] In some possible implementations, the server may convert the account content features of the sixth account and the account content features of the seventh account into the embedding vector of the sixth account and the embedding vector of the seventh account, respectively; and calculate the similarity between the embedding vector of the sixth account and the embedding vector of the seventh account to obtain the similarity between any two account clusters.

[0183] It should be understood that, in this implementable manner, the account content features of the sixth account may be R account content features, and similarly, the account content features of the seventh account may also be R account content features, where R is a positive integer.

[0184] In some implementations, the R account content features include at least one of the following: account name, account signature, content title, content frame extraction result, entity recognition result, OCR result, ASR result.

[0185] It should be understood that the explanation of the content features of the R accounts can be found above, and the embodiments of this application will not be repeated here.

[0186] In some implementations, if R is equal to 1, the server may convert the account content features of the sixth account into a corresponding embedding vector, and convert the account content features of the seventh account into a corresponding embedding vector.

[0187] In some implementations, if R is greater than 1, the server may convert multiple account content features of the sixth account into corresponding embedding vectors and concatenate these embedding vectors to form the embedding vector for the sixth account. Similarly, the server may convert multiple account content features of the seventh account into corresponding embedding vectors and concatenate these embedding vectors to form the embedding vector for the seventh account.

[0188] In this embodiment of the present application, the similarity can be cosine similarity, Euclidean distance, Manhattan distance, Pearson correlation coefficient or Jaccard similarity coefficient, etc., and this embodiment of the present application does not limit this.

[0189] A second implementation method is that the server can select a ninth account with the latest content publishing time from the first account cluster, and select a tenth account with the latest content publishing time from the second account cluster; based on the account content characteristics of the ninth account and the account content characteristics of the tenth account, determine the similarity between the first account cluster and the second account cluster.

[0190] It should be understood that the explanation of the content release time can be found above, and the embodiments of the present application will not be repeated here.

[0191] It should be understood that how to determine the similarity between the first account cluster and the second account cluster based on the account content characteristics of the ninth account and the account content characteristics of the tenth account can refer to the account content characteristics of the sixth account and the account content characteristics of the seventh account to determine the similarity between the first account cluster and the second account cluster, and the embodiments of the present application will not go into details about this.

[0192] A third implementation method is that the server can convert the account content features of each account in the first account cluster into an embedding vector corresponding to the account, and splice the embedding vectors corresponding to all accounts in the first account cluster to obtain the embedding vector corresponding to the first account cluster; similarly, the server can convert the account content features of each account in the second account cluster into an embedding vector corresponding to the account, and splice the embedding vectors corresponding to all accounts in the second account cluster to obtain the embedding vector corresponding to the second account cluster; further, the embedding vector corresponding to the first account cluster and the embedding vector corresponding to the second account cluster can be used to determine the similarity between the first account cluster and the second account cluster.

[0193] It should be understood that for each account in the first account cluster and the second account cluster, the number of account content features of the account may be R.

[0194] It should be understood that the explanation of the content features of the R accounts can be found above, and the embodiments of this application will not be repeated here.

[0195] In some implementations, if R is equal to 1, then for each account in the first account cluster and the second account cluster, the server may convert the account content feature of the account into a corresponding embedding vector.

[0196] In some implementations, if R is greater than 1, then for each account in the first account cluster and the second account cluster, the server may convert multiple account content features of the account into corresponding embedding vectors respectively, and concatenate these embedding vectors to form an embedding vector for the account.

[0197] In an embodiment of the present application, for any two account clusters, the server first calculates the similarity between the two account clusters. If the similarity between the two account clusters is greater than the preset similarity, the server can prompt the R&D personnel to determine whether to merge the two account clusters; since each account cluster includes large accounts, the merger between account clusters includes the merger between large accounts, and the merger between large accounts should be stricter to prevent large accounts that do not actually belong to the same operating entity from being mistakenly identified as belonging to the same operating entity.

[0198] As described above, the target model can calculate the scores of the two accounts based on their account content features. This score is used to determine whether the two accounts belong to the same operating entity. The following is a detailed explanation of the model training process:

[0199] Figure 4 A flow chart of a model training method provided in an embodiment of the present application is as follows: Figure 4 As shown, the method can be executed by a training device, which can be a server, which can be Figure 2The server 220 in the embodiment may also be other devices. The embodiment of the present application does not limit the training device. Figure 4 As shown, the method may include:

[0200] S410: Obtain an account pair and a label corresponding to the account pair for measuring whether two accounts in the account pair belong to the same operating entity;

[0201] There can be multiple account pairs, and each account pair and its corresponding tag can be obtained by any of the following feasible methods:

[0202] In some possible implementations, for each account, the training device can recall accounts similar to the account based on P account features of the account and P account features of other accounts other than the account that has published content, and then receive an object labeling operation, which is used to label whether the two accounts belong to the same operating entity. Furthermore, the training device generates a label corresponding to the account pair based on the object labeling operation.

[0203] It should be understood that the explanation of the characteristics of P accounts can be found above, and this embodiment of the present application will not be repeated here.

[0204] It should be understood that the method of recalling accounts during the training process can refer to the rough recall method mentioned above, and this embodiment of the present application will not be described in detail.

[0205] In some implementations, the value of the tag corresponding to the account pair can be 0 or 1. If the value of the tag is 0, it means that the account pair does not belong to the same operating entity. If the value of the tag is 1, it means that the account pair belongs to the same operating entity.

[0206] In other possible implementations, the training device can obtain failure cases (bad cases) of the product, which can be feedback from the object or discovered by the product itself. The product refers to a product that applies the content recommendation method provided in the embodiment of the present application. Each failure case includes: an account pair and a corresponding label for measuring whether the two accounts in the account pair belong to the same operating entity. For example, a failure case is <Account 1, Account 2, 1>, where 1 indicates that Account 1 and Account 2 belong to the same operating entity.

[0207] Through the above two feasible methods, approximately 100,000 account pairs and corresponding tags can be obtained.

[0208] S420: Obtain N account content features for each account in the account pair;

[0209] It should be understood that the explanation of the content features of N accounts can be found above, and the embodiments of this application will not be repeated here.

[0210] S430: N account content features of each account in the account pair and the label corresponding to the account pair constitute a training sample;

[0211] For example, Figure 5 A schematic diagram of account content features corresponding to an account pair provided in an embodiment of the present application, such as Figure 5 As shown in the figure, for <account 1, account 2, 1>, the account content features corresponding to account 1 include: account name, account signature, content title, content frame extraction results, entity recognition results, OCR results, and ASR results. The account content features corresponding to account 2 also include: account name, account signature, content title, content frame extraction results, entity recognition results, OCR results, and ASR results. Label 1 indicates that account 1 and account 2 belong to the same operating entity. The training sample corresponding to <account 1, account 2, 1> is <account content features of account 1, the aforementioned account content features of account 2, 1>.

[0212] S440: Train a target model based on the training samples.

[0213] The training device can input the account content features of each account pair in each training sample into the target model to obtain a prediction result, specifically a score, as to whether the two accounts in the account pair belong to the same operating entity. Furthermore, the training device can calculate the loss of the target model based on the prediction result and the corresponding label, and adjust the parameters of the target model based on the loss.

[0214] In some implementations, when the training device trains a target model based on training samples, the training device may perform modal fusion on N account content features, wherein the modal fusion may be performed through a transformation module (transformer) in the target model, but is not limited thereto.

[0215] In some implementations, when training the target model, the loss function used by the training device may be any of the following, but not limited to: L1 loss function, mean squared error (MSE) loss function, cross entropy loss function, etc.

[0216] In some implementations, after the target model is trained, it can be exported to ONNX format for use in subsequent execution stages.

[0217] In an embodiment of the present application, the training device may form a training sample from N account content features of each account in an account pair and the labels corresponding to the account pair; the target model is trained based on the training samples. Since the account content features can represent the account content, the target model can learn the account content features, so that in the subsequent execution stage, the target model can use the account content features to more accurately calculate the scores of the two accounts, thereby ensuring the accuracy of the content recommendation process.

[0218] The preferred embodiments of the present application are described in detail above in conjunction with the accompanying drawings. However, the present application is not limited to the specific details in the above embodiments. Within the technical concept of the present application, a variety of simple modifications can be made to the technical solution of the present application, and these simple modifications all fall within the scope of protection of the present application. For example, the various specific technical features described in the above specific embodiments can be combined in any suitable manner unless there is any contradiction. In order to avoid unnecessary repetition, the present application will not further explain various possible combinations. For another example, the various different embodiments of the present application can also be arbitrarily combined, and as long as they do not violate the ideas of the present application, they should also be regarded as the contents disclosed in the present application.

[0219] It should also be understood that in the various method embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0220] The above describes the method provided in the embodiment of the present application. The following describes the content recommendation device provided in the embodiment of the present application.

[0221] Figure 6 A schematic diagram of a content recommendation device 600 provided in an embodiment of the present application is shown as follows: Figure 6As shown, the device 600 includes: a determination module 610, a recall module 620, an input module 630 and a recommendation module 640; wherein the determination module 610 is used to determine the first account corresponding to the current recommended content of the target object; the determination module 610 is used to recall M second accounts similar to the first account from other accounts based on P account content features of the first account and P account content features of accounts other than the first account among the accounts that have published content; wherein P and M are both positive integers; the input module 630 is used to input N account content features of the first account and N account content features of the second account into the target model for each second account in the M second accounts to obtain scores for the first account and the second account; wherein the scores are used to measure whether the first account and the second account belong to the same operating entity; N is a positive integer; the determination module 610 is further used to determine a target account from the M second accounts that belongs to the same operating entity as the first account based on the respective scores of the first account and the M second accounts; and the recommendation module 640 is used to prohibit recommending the content of the target account to the target object within a preset time period.

[0222] In some implementations, the determination module 610 is specifically configured to: determine K third accounts with scores greater than a target preset score among the M second accounts; where K is a positive integer; and determine a target account based on the K third accounts.

[0223] In some implementations, the determination module 610 is specifically used to: determine L fourth accounts among the K third accounts whose follower counts are greater than a first preset number; where L is a positive integer; and determine a target account based on the L fourth accounts.

[0224] In some implementations, the determination module 610 is specifically used to: select the fifth account with the latest content release time from the L fourth accounts; classify the first account into the account cluster to which the fifth account belongs; and determine the account in the account cluster to which the fifth account belongs as the target account.

[0225] In some implementations, the determination module 610 is further configured to: for each of the L fourth accounts, determine the publishing time of the latest content published by the fourth account as the content publishing time of the fourth account.

[0226] In some implementations, the determination module 610 is further configured to: for each of the L fourth accounts, determine the average publishing time of the fourth account's most recent Q published contents as the content publishing time of the fourth account; wherein Q is an integer greater than 1.

[0227] In some possible implementations, the device 600 also includes: a sampling module 650, an acquisition module 660, and an adjustment module 670, wherein before the determination module 610 determines K third accounts with scores greater than the target preset scores among the M second accounts based on the scores of the first account and the M second accounts, the sampling module 650 is used to sample the M second accounts to obtain multiple sampled accounts; the determination module 610 is also used to determine the machine judgment result of whether the first account and the multiple sampled accounts belong to the same operating entity based on the scores of the first account and the multiple sampled accounts and the initial preset scores; the acquisition module 660 is used to obtain the manual judgment result of whether the first account and the multiple sampled accounts belong to the same operating entity; the adjustment module 670 is used to adjust the initial preset score based on the machine judgment result and the manual judgment result to obtain the target preset score.

[0228] In some possible implementations, the device 600 also includes: a training module 680, wherein the acquisition module 660 is further used to: obtain account pairs and labels corresponding to the account pairs for measuring whether the two accounts in the account pairs belong to the same operating entity; obtain N account content features of each account in the account pair; the determination module 610 is further used to form a training sample from the N account content features of each account in the account pair and the labels corresponding to the account pairs; the training module 680 is used to train the target model based on the training samples.

[0229] In some possible implementations, the recall module 620 is specifically used to: if P is greater than or equal to 2, then for each account content feature in the P account content features, based on the account content feature of the first account and the account content features of the other accounts, recall a second account similar to the first account from other accounts.

[0230] In some possible implementations, the device 600 also includes: a push module 690, wherein the determination module 610 is further used to determine the similarity between the first account cluster and the second account cluster; wherein the first account cluster and the second account cluster are any two account clusters; the push module 690 is used to: if the similarity between the first account cluster and the second account cluster is greater than a preset similarity, then push a prompt message to prompt the R&D personnel to determine whether to merge the first account cluster and the second account cluster.

[0231] In some possible implementations, the determination module 610 is specifically used to: select a sixth account in the first account cluster whose number of followers is greater than a second preset number, and select a seventh account in the second account cluster whose number of followers is greater than a third preset number; and determine the similarity between the first account cluster and the second account cluster based on the account content characteristics of the sixth account and the account content characteristics of the seventh account.

[0232] In some implementations, the determination module 610 is specifically configured to: if the first account cluster includes multiple first candidate accounts whose follower counts are greater than a second preset number, select the first candidate account with the latest content publishing time from the multiple first candidate accounts as the sixth account.

[0233] In some implementations, the determination module 610 is specifically configured to: if the second account cluster includes multiple second candidate accounts whose follower counts are greater than a third preset number, select the second candidate account with the latest content publishing time from the multiple second candidate accounts as the seventh account.

[0234] In some implementations, the determination module 610 is specifically used to: convert the account content features of the sixth account and the account content features of the seventh account into the embedding vector of the sixth account and the embedding vector of the seventh account, respectively; calculate the similarity between the embedding vector of the sixth account and the embedding vector of the seventh account to obtain the similarity between any two account clusters.

[0235] In some implementations, the P account content features include at least one of the following: account name, content extraction result.

[0236] In some implementations, the N account content features include at least one of the following: account name, account signature, content title, content frame extraction result, entity recognition result, OCR result, ASR result.

[0237] It should be understood that the device embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, they will not be described here. Specifically, Figure 6 The apparatus 600 shown may perform Figure 2 as well as Figure 4 The corresponding method embodiment, and the aforementioned and other operations and / or functions of each module in the device 600 are respectively to achieve Figure 2 as well as Figure 4 For the sake of brevity, the corresponding processes in each method are not repeated here.

[0238] The above describes the device 600 of the embodiment of the present application from the perspective of functional modules in conjunction with the accompanying drawings. It should be understood that the functional module can be implemented in hardware form, can be implemented by instructions in software form, and can also be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software form instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps in the above method embodiment in conjunction with its hardware.

[0239] Figure 7 It is a schematic block diagram of an electronic device provided in an embodiment of the present application.

[0240] like Figure 7 As shown, the electronic device may include:

[0241] The memory 710 and the processor 720 are configured to store computer programs and transmit the program code to the processor 720. In other words, the processor 720 can call and run the computer program from the memory 710 to implement the method in the embodiment of the present application.

[0242] For example, the processor 720 may be configured to execute the above method embodiments according to instructions in the computer program.

[0243] In some embodiments of the present application, the processor 720 may include but is not limited to:

[0244] General-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc.

[0245] In some embodiments of the present application, the memory 710 includes but is not limited to:

[0246] Volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DR RAM).

[0247] In some embodiments of the present application, the computer program may be divided into one or more modules, which are stored in the memory 710 and executed by the processor 720 to implement the method provided by the present application. The one or more modules may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.

[0248] like Figure 7 As shown, the electronic device may further include:

[0249] The transceiver 730 may be connected to the processor 720 or the memory 710 .

[0250] The processor 720 may control the transceiver 730 to communicate with other devices. Specifically, the processor 720 may send information or data to other devices or receive information or data sent by other devices. The transceiver 730 may include a transmitter and a receiver. The transceiver 730 may further include one or more antennas.

[0251] It should be understood that the various components in the electronic device are connected via a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus and a status signal bus.

[0252] The present application also provides a computer storage medium having a computer program stored thereon, which, when executed by a computer, enables the computer to perform the method of the above-mentioned method embodiment. In other words, the present application also provides a computer program product containing instructions, which, when executed by a computer, enables the computer to perform the method of the above-mentioned method embodiment.

[0253] When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a digital video disc (DVD)), or a semiconductor medium (e.g., a solid state drive (SSD)).

[0254] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0255] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0256] Modules described as separate components may or may not be physically separate, and components displayed as modules may or may not be physical modules, i.e., they may be located in one place or distributed across multiple network elements. Some or all of the modules may be selected based on actual needs to achieve the purpose of the present embodiment. For example, the functional modules in the various embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module.

[0257] The above content is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A content recommendation method, characterized in that: include: Determine the first account corresponding to the current recommended content of the target object; Based on P account content features of the first account and P account content features of accounts other than the first account that have published content, recall M second accounts similar to the first account from the other accounts; where P and M are both positive integers; For each of the M second accounts, input the N account content features of the first account and the N account content features of the second account into the target model to obtain scores for the first account and the second account; wherein the scores are used to measure whether the first account and the second account belong to the same operating entity; N is a positive integer; Based on the scores of the first account and the M second accounts, determining a target account from the M second accounts that belongs to the same operating entity as the first account; It is prohibited to recommend the content of the target account to the target object within a preset time period.

2. The method according to claim 1, characterized in that The determining, based on the scores of the first account and the M second accounts, a target account belonging to the same operating entity as the first account from the M second accounts includes: Determining K third accounts from the M second accounts whose scores are greater than a target preset score; wherein K is a positive integer; The target account is determined based on the K third accounts.

3. The method according to claim 2, characterized in that The determining the target account based on the K third accounts includes: Determining L fourth accounts from the K third accounts, each of which has a following amount greater than a first preset number; where L is a positive integer; The target account is determined based on the L fourth accounts.

4. The method according to claim 3, characterized in that The determining the target account based on the L fourth accounts includes: Selecting a fifth account with the latest content publishing time from the L fourth accounts; assigning the first account to the account cluster to which the fifth account belongs; An account in the account cluster to which the fifth account belongs is determined as the target account.

5. The method according to claim 4, characterized in that Also includes: For each of the L fourth accounts, the publishing time of the latest content published by the fourth account is determined as the content publishing time of the fourth account.

6. The method according to claim 4, characterized in that Also includes: For each of the L fourth accounts, the average publishing time of the most recent Q content published by the fourth account is determined as the content publishing time of the fourth account; wherein Q is an integer greater than 1.

7. The method according to any one of claims 2 to 6, characterized in that: Before determining K third accounts having scores greater than a target preset score from the M second accounts based on the scores of the first account and the M second accounts, the method further includes: Sampling the M second accounts to obtain a plurality of sampled accounts; Determining, based on the scores of the first account and the multiple sampled accounts and the initial preset scores, a machine-generated result of determining whether the first account and the multiple sampled accounts belong to the same operating entity; Obtaining manual determination results of whether the first account and the plurality of sampled accounts belong to the same operating entity; The initial preset score is adjusted based on the machine discrimination result and the manual discrimination result to obtain the target preset score.

8. The method according to any one of claims 1 to 6, characterized in that Also includes: Obtain an account pair and a label corresponding to the account pair for measuring whether the two accounts in the account pair belong to the same operating entity; Obtain N account content features for each account in the account pair; The N account content features of each account in the account pair and the labels corresponding to the account pair constitute a training sample; The target model is trained based on the training samples.

9. The method according to any one of claims 1 to 6, characterized in that The step of recalling M second accounts similar to the first account from the other accounts based on the P account content features of the first account and the P account content features of the other accounts includes: If P is greater than or equal to 2, for each of the P account content features, based on the account content feature of the first account and the account content features of the other accounts, a second account similar to the first account is recalled from the other accounts.

10. The method according to any one of claims 1 to 6, characterized in that Also includes: Determining a similarity between a first account cluster and a second account cluster; wherein the first account cluster and the second account cluster are any two account clusters; If the similarity between the first account cluster and the second account cluster is greater than a preset similarity, a prompt message is pushed to prompt the R&D personnel to determine whether to merge the first account cluster and the second account cluster.

11. The method according to claim 10, characterized in that Determining the similarity between the first account cluster and the second account cluster includes: Selecting a sixth account in the first account cluster whose follower count is greater than a second preset count, and selecting a seventh account in the second account cluster whose follower count is greater than a third preset count; Based on the account content feature of the sixth account and the account content feature of the seventh account, the similarity between the first account cluster and the second account cluster is determined.

12. The method according to claim 11, characterized in that The selecting a sixth account in the first account cluster whose following amount is greater than a second preset amount includes: If the first account cluster includes multiple first candidate accounts whose follower counts are greater than the second preset number, the first candidate account with the latest content publishing time is selected from the multiple first candidate accounts as the sixth account.

13. The method according to claim 11, characterized in that The selecting a seventh account in the second account cluster whose following amount is greater than a third preset number includes: If the second account cluster includes multiple second candidate accounts whose follower counts are greater than the third preset number, the second candidate account with the latest content publishing time is selected from the multiple second candidate accounts as the seventh account.

14. The method according to claim 11, characterized in that The determining the similarity between the arbitrary two account clusters based on the account content feature of the sixth account and the account content feature of the seventh account includes: Converting the account content features of the sixth account and the account content features of the seventh account into an embedding vector of the sixth account and an embedding vector of the seventh account, respectively; The similarity between the embedding vector of the sixth account and the embedding vector of the seventh account is calculated to obtain the similarity between the arbitrary two account clusters.

15. The method according to any one of claims 1 to 6, characterized in that The P account content features include at least one of the following: account name, content frame extraction result.

16. The method according to any one of claims 1 to 6, characterized in that The N account content features include at least one of the following: account name, account signature, content title, content frame extraction result, entity recognition result, optical character recognition (OCR) result, and automatic speech recognition (ASR) result.

17. A content recommendation device, characterized in that: include: Determination module, recall module, input module and recommendation module; The determination module is used to determine the first account corresponding to the current recommended content of the target object; The recall module is configured to recall M second accounts similar to the first account from among the accounts that have published content, based on P account content features of the first account and P account content features of accounts other than the first account that have published content; wherein P and M are both positive integers; The input module is configured to input, for each of the M second accounts, N account content features of the first account and N account content features of the second account into a target model to obtain scores for the first account and the second account; wherein the scores are used to measure whether the first account and the second account belong to the same operating entity; N is a positive integer; The determination module is further configured to determine, based on the scores of the first account and the M second accounts, a target account belonging to the same operating entity as the first account from among the M second accounts; The recommendation module is used to prohibit recommending the content of the target account to the target object within a preset time period.

18. An electronic device, characterized in that: include: A processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory to execute the method according to any one of claims 1 to 16.

19. A computer-readable storage medium, characterized in that Used to store a computer program, wherein the computer program causes a computer to execute the method according to any one of claims 1 to 16.

20. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 16 is implemented.