Training evaluation method, device, equipment and storage medium based on recall model

By evaluating the account characteristics and video characteristics of the recall model offline, computed the matching degree and recall rate, the problem of small resources and scale occupied by online AB tests in the existing technology is solved, and efficient and timely model evaluation and parameter iteration are achieved.

CN114357242BActive Publication Date: 2025-05-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111575932.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-20
Publication Date
2025-05-06
Estimated Expiration
2041-12-20

AI Technical Summary

Technical Problem

The evaluation method of the existing recall model mainly relies on online AB testing, occupying online resources, and has a small test scale and long observation period, which cannot effectively improve the model training effect.

Method used

By obtaining the account features and video features extracted during the online training of the recall model, it is stored in the feature library, and performing offline evaluation, calculating the matching degree between the target account features and the video features, sorting and selecting the specified ranking, and calculating the recall rate based on the number of positive samples to evaluate the training effect of the model.

Benefits of technology

This method allows the evaluation of the recall model to not occupy online resources, saves online machine resources, improves the timeliness of evaluation, improves the efficiency of model parameter iteration, reduces evaluation deviations, and improves evaluation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114357242B_ABST
    Figure CN114357242B_ABST
Patent Text Reader

Abstract

The embodiments of the present application disclose a training evaluation method and device based on a recall model, an electronic device, and a storage medium, which can be applied to the fields of autonomous driving, smart transportation, etc., including: storing the account features and video features extracted based on the training video samples during the online training of the recall model in the account feature library and the video feature library respectively; sampling the target account features from the account feature library offline, searching the video feature set related to the entry time from the video feature library for the target account features; calculating the matching degree between the target account features and each target video feature in the video feature set, and selecting a specified ranking based on the sorting of the matching degree values ​​from large to small; calculating the recall rate based on the number of positive samples associated with the target video features corresponding to the specified ranking, and the positive sample data associated with the target video features in the video feature set. The scheme of the embodiments of the present application can save online machine resources and improve the iteration effect of model parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a training evaluation method and device based on a recall model, an electronic device, a storage medium, and a program product. Background Art

[0002] Recommendation system refers to a system in the Internet era where a platform automatically selects / matches products on the platform based on the interests of the service recipient and presents them to the service recipient. Recommendation system usually includes a recall model, which is used to select a subset from the candidate pool that meets the target and computing power constraints.

[0003] In order to improve the training effect of the recall model, the recall model needs to be evaluated. At present, the recall model is usually evaluated by online AB testing, but online AB testing requires online resources, and the scale of online testing is small and the observation period is long. Summary of the invention

[0004] To solve the above technical problems, the embodiments of the present application provide a training evaluation method and device based on a recall model, an electronic device, a storage medium, and a program product.

[0005] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.

[0006] According to one aspect of an embodiment of the present application, a training evaluation method based on a recall model is provided, comprising:

[0007] Acquire the account features and video features extracted based on the training video samples during the online training of the recall model, and store the account features and the video features in an account feature library and a video feature library respectively;

[0008] Offline sampling is performed from the account feature library to obtain target account features, and a video feature set related to the storage time is searched from the video feature library for the target account features;

[0009] Calculating the matching degree between the target account feature and each target video feature contained in the video feature set respectively, and selecting a designated ranking based on the sorting of the matching degree values ​​from large to small;

[0010] A recall rate is calculated according to the number of positive samples associated with the target video feature corresponding to the specified ranking and the positive sample data associated with the target video feature in the video feature set, and the recall rate is used to evaluate the training effect of the recall model.

[0011] According to one aspect of an embodiment of the present application, a training evaluation device based on a recall model is provided, comprising:

[0012] A feature acquisition module, configured to acquire account features and video features extracted based on training video samples during the online training of the recall model, and store the account features and video features in an account feature library and a video feature library, respectively;

[0013] An offline sampling module, configured to obtain target account features from the account feature library by offline sampling, and search the video feature set related to the storage time from the video feature library for the target account features;

[0014] A matching ranking module, configured to respectively calculate the matching degree between the target account feature and each target video feature contained in the video feature set, and select a designated ranking based on the order of matching degree values ​​from large to small;

[0015] The recall rate calculation module is configured to calculate the recall rate based on the number of positive samples associated with the target video feature corresponding to the specified ranking and the positive sample data associated with the target video feature in the video feature set, and the recall rate is used to evaluate the training effect of the recall model.

[0016] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the electronic device implements the recall model-based training evaluation method as described above.

[0017] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor of an electronic device, the electronic device executes the recall model-based training evaluation method as described above.

[0018] According to one aspect of an embodiment of the present application, a computer program product is provided, including a computer program, which, when executed by a processor, implements the training evaluation method based on the recall model as described above.

[0019] In the technical solution provided in the embodiments of the present application, the recall model is obtained based on the account features and video features extracted from the training video samples during the online training process, and the recall model is evaluated based on the acquired account features and video features, so that the evaluation process does not occupy online resources and saves online machine resources; and, evaluating the recall model during the training process can improve the timeliness of the evaluation and improve the efficiency of model parameter iteration.

[0020] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0022] Figure 1 It is a schematic diagram of an implementation environment involved in this application;

[0023] Figure 2 is a flowchart of a training evaluation method based on a recall model shown in an exemplary embodiment of the present application;

[0024] Figure 3 yes Figure 2 A flowchart of step S130 in the illustrated embodiment in an exemplary embodiment;

[0025] Figure 4 is a flowchart of a training evaluation method based on a recall model shown in another exemplary embodiment of the present application;

[0026] Figure 5 is a flowchart of training and evaluating a recall model shown in an exemplary embodiment of the present application;

[0027] Figure 6 is a structural schematic diagram of a training evaluation device based on a recall model shown in an exemplary embodiment of the present application;

[0028] Figure 7 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application is shown. DETAILED DESCRIPTION

[0029] Here, exemplary embodiments will be described in detail, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are only examples of devices and methods consistent with some aspects of the present application as detailed in the attached claims.

[0030] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0031] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0032] It should also be noted that the "multiple" mentioned in this application refers to two or more than two. "And / or" describes the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship.

[0033] Before introducing the technical solutions of the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are explained first. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0034] Recommendation system: refers to a system in the Internet era where a platform automatically selects / matches products on the platform based on the interests of the service recipients and presents them to the service recipients. Since the service recipients' recent behaviors express stronger interests or trends, in order to improve real-time performance, the current recommendation system widely uses real-time technology, that is, through streaming data transmission / model training at the minute level, export and launch / new model update at the minute level and online reasoning, it can respond to and reason about the behavior of individual and group service recipients in a timely manner. Due to the limitations of the recommendation system's computing power and online system latency, the recommendation system can adopt a funnel-level structure of recall-coarse sorting (optional)-fine sorting-strategy (mixed sorting).

[0035] Recall: Used to select a subset from the entire candidate pool that meets the target and computing power limit.

[0036] Online AB testing: Before a new function is fully launched, the online traffic is split and a small portion of the split traffic is used to test the new function and evaluate its effectiveness.

[0037] Cloud storage is a new concept that extends and develops from the concept of cloud computing. A distributed cloud storage system refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to bring together a large number of different types of storage devices (storage devices are also called storage nodes) in the network through application software or application interfaces to work together and provide external data storage and business access functions.

[0038] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0039] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operating / interactive systems, mechatronics and other technologies. Artificial intelligence software technologies mainly include computer vision technology, speech processing technology, natural language processing technology, as well as machine learning / deep learning, autonomous driving, smart transportation and other major directions.

[0040] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.

[0041] At present, the storage method of the distributed cloud storage system is: create a logical volume, and when creating a logical volume, allocate physical storage space for each logical volume. The physical storage space may be composed of disks of a storage device or several storage devices. The client stores data on a logical volume, that is, stores the data on the file system. The file system divides the data into many parts, each of which is an object. The object contains not only data but also additional information such as data identification (ID, IDentity). The file system writes each object to the physical storage space of the logical volume, and the file system records the storage location information of each object, so that when the client requests to access the data, the file system can allow the client to access the data according to the storage location information of each object.

[0042] The process of allocating physical storage space to logical volumes in a distributed cloud storage system is as follows: based on the estimated capacity of objects stored in the logical volumes (this estimate often has a large margin relative to the actual capacity of the objects to be stored) and the grouping of independent redundant disk arrays (RAID, Redundant Array of Independent Disks), the physical storage space is divided into stripes in advance. A logical volume can be understood as a stripe, thereby allocating physical storage space to the logical volume.

[0043] In order to improve the training effect of the recall model, the recall model needs to be evaluated. At present, the recall model is usually evaluated by online AB testing. However, during the online AB testing process, an online environment is required, which occupies online machine resources, such as memory, external storage and other resources; and the traffic used for online testing is generally less than 20%, which is relatively small in scale. Based on this, the embodiments of the present application provide a training evaluation method and device based on a recall model, an electronic device, a storage medium, and a program product, so that the evaluation process of the recall model does not occupy online resources, and the evaluation scale can be adjusted by using adaptive adjustment.

[0044] See also Figure 1 , Figure 1 Schematic diagram of an implementation environment involved in the present application. The implementation environment includes a training evaluation device 100 based on a recall model, a recall model 200 and an online training device 300.

[0045] The training evaluation device 100 may be a server or other device. The server may be a server that provides various services, which may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, etc. This is not limited here.

[0046] The online training device 300 may also be a server or other equipment.

[0047] During the online training process of the recall model 200 by the online training device 300, the recall model will extract account features and video features based on the training video samples. The training evaluation device 100 can obtain the account features and video features extracted by the recall model based on the training video samples during the online training process, and store the account features and video features in the account feature library and the video feature library respectively; obtain the target account features from the account feature library by offline sampling, and search the video feature set related to the storage time from the video feature library for the target account features; respectively calculate the matching degree between the target account features and each target video feature contained in the video feature set, and select the designated ranking based on the sorting of the matching degree values ​​from large to small; calculate the recall rate according to the number of positive samples associated with the target video features corresponding to the designated ranking, and the positive sample data associated with the target video features in the video feature set, and the recall rate is used to evaluate the training effect of the recall model. In this way, by obtaining the account features and video features extracted from the training video samples during the online training of the recall model, and evaluating the recall model based on the acquired account features and video features, the evaluation process does not occupy online resources, saving online machine resources; and, evaluating the recall model during the training process can improve the timeliness of the evaluation and the efficiency of the model parameter iteration. In addition, the data used in the evaluation process is the same as the data used in the training process, thereby reducing deviations and improving evaluation accuracy.

[0048] The account model is a model created based on machine learning; the account feature library and the video feature library can be stored in a storage system, which can be a storage system based on cloud storage technology, or other types of storage systems. Figure 2 , Figure 2 is a flowchart of a training evaluation method based on a recall model shown in an exemplary embodiment of the present application. The method can be applied to Figure 1 The implementation environment shown, which can be Figure 1The recall model-based training evaluation apparatus 100 is executed in the illustrated implementation environment.

[0049] After the recall model meets certain conditions, the recall model can be put online to provide services to terminals, where the terminals include but are not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, car terminals, etc.

[0050] It should be noted that in addition to the aforementioned application scenarios, the embodiments of the present application can also be applied to various application scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, etc. In practical applications, corresponding adjustments can be made according to specific application scenarios. For example, if applied to cloud technology scenarios, the recall model can be deployed in the cloud, and the account feature library and video feature library can also be stored based on cloud storage technology; if applied to artificial intelligence or assisted driving scenarios, the recall model can be deployed in vehicle terminals, navigation terminals, etc., for navigation, assisted driving, etc.

[0051] like Figure 2 As shown, in an exemplary embodiment, the training evaluation method based on the recall model may include steps S110 to S140, which are described in detail as follows:

[0052] Step S110, obtaining the account features and video features extracted based on the training video samples during the online training of the recall model, and storing the account features and video features in the account feature library and the video feature library respectively.

[0053] It should be noted that in this embodiment, the recall model is applied to the video recall scenario to select a model that meets the target and computing power from the candidate pool. The recall model can be a model created based on DNN (Deep Neural Networks), or of course, a model created based on other machine learning networks. The candidate pool can be flexibly set according to actual needs, for example, including but not limited to the videos contained in the video platform.

[0054] The training video sample is a video used to train the recall model, which may include positive samples and negative samples. The training video sample may be determined based on online real-time messages. For example, during the operation of the video platform, real-time messages will be generated, and account data may be obtained from the real-time messages, including but not limited to account attribute information, account behavior information, etc., wherein the account attribute information includes but is not limited to the age, gender, etc. of the service object corresponding to the account, and the account behavior information includes but is not limited to clicks, views, likes, comments, forwarding, and other behaviors. Based on the account data, the account characteristics of the account may be determined, and positive samples and negative samples may be constructed based on the account characteristics, wherein the positive sample may include videos that the account has watched for a certain time (for example, 10 minutes, 3 minutes, etc.), videos forwarded by the account, videos liked by the account, etc., and the negative sample may include videos that the account has not watched, videos that the account has blocked, etc., or videos may be randomly selected from the candidate pool as negative samples.

[0055] The account feature library is used to store account features, and its specific type can be flexibly set according to actual needs. For example, in one example, since the account features are updated in real time, the types of the account feature library include but are not limited to real-time tables. Real-time tables refer to table files whose contents are updated in real time to meet the needs of account features extracted in real time by the recall model during online training.

[0056] The video feature library is used to store the video features of the videos in the candidate pool, and its specific type can be flexibly set according to actual needs. In one example, since the amount of videos in the candidate pool is large, usually in the millions to billions, in order to reduce storage pressure, the types of video feature libraries include but are not limited to distributed file systems. Among them, a distributed file system (Distributed File System, DFS) refers to a physical storage resource managed by a file system that is not necessarily directly connected to a local node, but is connected to a node through a computer network, or is a complete hierarchical file system formed by combining several different logical disk partitions or volumes, which can not only reduce single-point storage pressure, but also meet the video feature storage requirements based on time points.

[0057] When the recall model is trained online, the recall model extracts account features based on the training video samples and extracts features from the videos in the candidate pool to obtain video features for each video in the candidate pool.

[0058] During the online training process of the recall model, the input training video samples associated with the account can be identified to extract the account features of the account. In addition, the recall model will also identify the videos in the candidate pool and extract the video features of each video, so as to facilitate the subsequent recall of the video corresponding to the account features based on the account features and the video features. In order to evaluate the recall model, in this embodiment, the account features and video features extracted by the recall model based on the training video samples during the online training process are obtained, and the account features are stored in the account feature library, and the video features are stored in the video feature library.

[0059] In some embodiments, the candidate pool is continuously updated, and there will be different versions of the candidate pool. The recall model can extract features from videos in different versions of the candidate pool. When recalling, it is usually to recall qualified videos from a certain version of the candidate pool. Therefore, when obtaining video features during the online training process of the recall model, the video features corresponding to the videos contained in the cache pool can be obtained in units of the cache pool.

[0060] Step S120, sampling the target account features from the account feature library offline, and searching the video feature set related to the storage time from the video feature library for the target account features.

[0061] The account feature library stores account features of different accounts acquired at different times, and the amount of data is large. Therefore, in this embodiment, offline sampling can be performed in the account feature library to obtain the target account features. The offline sampling method can be flexibly set according to actual needs.

[0062] The videos in the candidate pool are constantly updated. Therefore, there are different versions of the candidate pool. The recall model will extract features from the videos in different versions of the candidate pool. Correspondingly, the video feature library stores video features corresponding to different versions of the candidate pool. For example, at a certain moment, the data in candidate pool 1 is updated, and an updated candidate pool is obtained, recorded as candidate pool 2. The video feature library stores the video features corresponding to candidate pool 1 and the video features corresponding to candidate pool 2. The video features corresponding to the candidate pool are the video features corresponding to the videos contained in the candidate pool.

[0063] The storage time of the video feature may be the time when the video feature is stored in the video feature library, or the time when the video feature is obtained; the storage time of the account feature may be the time when the account feature is stored in the account feature library, or the time when the account feature is obtained;

[0064] In this embodiment, after obtaining the target account feature from the account feature library through offline sampling, the video feature corresponding to the storage time of the target account feature is obtained from the video feature library to obtain a video feature set. For example, the video feature whose storage time is earlier than the storage time of the target account feature can be obtained.

[0065] Step S130, respectively calculating the matching degree between the target account feature and each target video feature contained in the video feature set, and selecting a designated ranking based on the sorting of the matching degree values ​​from large to small.

[0066] Among them, the designated ranking can be flexibly set according to actual needs, for example, the top 100, 500, etc.

[0067] In this embodiment, the matching degree between the target account feature and each target video feature contained in the video feature set is calculated respectively, and based on the value of the matching degree, the target video feature corresponding to the specified ranking is selected in descending order.

[0068] Step S140, calculating the recall rate according to the number of positive samples associated with the target video feature corresponding to the specified ranking and the positive sample data associated with the target video feature in the video feature set, and the recall rate is used to evaluate the training effect of the recall model.

[0069] The number of positive samples associated with the target video feature corresponding to the specified ranking includes: the number of positive samples in the video associated with the target video feature corresponding to the specified ranking. The positive sample data associated with the target video feature in the video feature set includes: the number of positive samples in the video associated with the target video feature in the video feature set.

[0070] The recall rate can be the ratio of the number of positive samples to the positive sample data. In one example, assuming that the video feature set includes the features of video a, video b, video c, video d, and video e, where the positive samples are video a and video c, the specified ranking is 3, and the order of the matching values ​​from large to small is the features of video a, the features of video b, the features of video e, the features of video c, and the features of video d. Since the top 3 only include the features of the positive sample video a, the number of positive samples associated with the target video features corresponding to the specified ranking is 1 (i.e., video a), and the positive sample data associated with the target video features in the video feature set is 2 (i.e., videos a and c), and the recall rate is 1 / 2=0.5.

[0071] Since both account features and video features are obtained from the online training process of the recall model, they can characterize the training effect of the recall model to a certain extent. Therefore, the recall rate calculated based on the number of positive samples associated with the target video features corresponding to the specified ranking and the positive sample data associated with the target video features in the video feature set can evaluate the training effect of the recall model.

[0072] In this embodiment, the account features and video features extracted based on the training video samples during the online training of the recall model are obtained, and the account features and video features are stored in the account feature library and the video feature library respectively; the target account features are obtained by offline sampling from the account feature library, and the video feature set related to the storage time is searched from the video feature library for the target account features; the matching degree between the target account features and each target video feature contained in the video feature set is calculated respectively, and a specified ranking is selected based on the sorting of the matching degree values ​​from large to small; the recall rate is calculated according to the number of positive samples associated with the target video features corresponding to the specified ranking, and the positive sample data associated with the target video features in the video feature set. The recall rate is used to evaluate the training effect of the recall model, so that the evaluation process does not occupy online resources and saves online machine resources; and, evaluating the recall model during the training process can improve the timeliness of the evaluation and the efficiency of model parameter iteration.

[0073] In an exemplary embodiment, since it is necessary to refer to the time parameter when determining the video feature set, the training evaluation method based on the recall model may further include: in the process of storing the account features and the video features in the account feature library and the video feature library respectively, the acquisition time of the account features and the acquisition time of the video features are also recorded in the account feature library and the video feature library respectively. This facilitates the subsequent search for the time-related video feature set from the video feature library for the target account features.

[0074] See also Figure 3 , Figure 3 Under the condition that the acquisition time of the account feature and the acquisition time of the video feature are recorded in the account feature library and the video feature library respectively, Figure 2 The flowchart of step S130 in the illustrated embodiment in an exemplary embodiment is as follows: Figure 3 As shown, the process of obtaining target account features by offline sampling from the account feature library and searching the video feature set related to the storage time from the video feature library for the target account features may include steps S131-S132, which are described in detail as follows:

[0075] Step S131, sampling target account features from the account feature library offline, and determining the acquisition time corresponding to the target account features.

[0076] In order to select a video feature set corresponding to the target account feature, in this embodiment, during the process of offline sampling of the target account feature from the account feature library, the acquisition time corresponding to the target account feature can be obtained.

[0077] Step S132, searching the video feature library for a video feature whose acquisition time is earlier than the acquisition time of the target account feature and closest to the acquisition time of the target account feature, and using the searched video feature as the target video feature in the video feature set.

[0078] Since the recall model can only recall videos from the candidate pool whose corresponding time is earlier than the generation time of the account feature when recalling videos from the candidate pool based on the account feature, in order to improve the accuracy of offline evaluation, in this embodiment, after sampling the target account feature offline from the account feature library and determining the acquisition time corresponding to the target account feature, the video feature whose acquisition time is earlier than the acquisition time of the target account feature and closest to the acquisition time of the target account feature is searched from the video feature library, and the searched video feature is used as the target video feature in the video feature set. For example, assuming that the acquisition time of the target account feature is 12:10:05, the video feature library includes the video features corresponding to candidate pool 1, the video features corresponding to candidate pool 2, and the video features corresponding to candidate pool 3. The acquisition time of the video features corresponding to candidate pool 1 is 12:3:06, the acquisition time of the video features corresponding to candidate pool 2 is 12:6:06, and the acquisition time of the video features corresponding to candidate pool 1 is 12:20:06. Therefore, the video features corresponding to candidate pool 2 are used as the target video features in the video feature set.

[0079] In some embodiments, if there are multiple target account features, considering that the acquisition time of each target account feature may be different, therefore, for each target account feature, the corresponding video feature set can be determined according to the acquisition time of the target account feature, and then, based on the number of positive samples included in the video associated with the corresponding video feature set and the data of positive samples included in the video associated with the target video feature of the specified ranking, the recall rate of the target video feature is calculated; then, the recall rates corresponding to the multiple target account features are averaged to obtain the final recall rate. Alternatively, if there are multiple target account features, the earliest acquisition time can be determined from the acquisition times corresponding to the multiple target account features, and a video feature set is determined based on the earliest acquisition time, and then, for each target account feature, the number of positive samples included in the video associated with it in the video feature set and the data of positive samples included in the video associated with the target video feature of the specified ranking are determined, and the recall rate of the target video feature is calculated; then, the recall rates corresponding to the multiple target account features are averaged to obtain the final recall rate. Of course, there can be other processing methods, which are not limited here.

[0080] In this embodiment, the target account features are sampled offline from the account feature library, and the acquisition time corresponding to the target account features is determined. The video feature library is searched for video features whose acquisition time is earlier than the acquisition time of the target account features and closest to the acquisition time of the target account features. The searched video features are used as the target video features in the video feature set, thereby improving the evaluation accuracy.

[0081] In an exemplary embodiment, Figure 2 In step S130 of the illustrated embodiment, the process of sampling target account features offline from the account feature library and determining the acquisition time corresponding to the target account features may include: periodically sampling a specified number of target account features offline from the account feature library based on a preset time interval, wherein the preset time interval is greater than the frequency of extracting account features and video features during the online training process of the recall model.

[0082] The preset time interval is the time interval for evaluating the recall model, that is, the target account features are obtained at a certain time interval, and the recall model is evaluated based on the target account features. The specific value of the preset time interval can be flexibly set according to actual needs, for example, it can be set to 1 hour, 2 hours, etc. The specified number is the number of target account features sampled in each time interval, and its specific value can be flexibly set according to actual needs, for example, it can be 100,000, 50,000, etc.

[0083] During the online training process, the recall model will extract account features from newly added training samples at regular intervals. Since the candidate pool is constantly updated, the recall model can also extract features from videos included in the new version of the candidate pool at regular intervals. For example, the recall module can extract account features at the minute level, and extract features from videos included in the candidate pool at the minute level. Minute-level extraction means that the interval between two extractions is at the minute level (less than 1 hour, for example, 1 minute, 2 minutes, etc.).

[0084] In order to avoid multiple evaluations of the recall model based on the same account features and video features, in this embodiment, the preset time interval is greater than the frequency of extracting account features and video features during the online training process of the recall model.

[0085] In this embodiment, a specified number of target account features are periodically sampled offline from the account feature library based on a preset time interval, so that the recall model can be evaluated once every preset time interval, thereby improving the real-time performance of the evaluation and facilitating comparison of recall models in different time periods.

[0086] In an exemplary embodiment, see Figure 4As shown, after calculating the recall rate according to the number of positive samples associated with the target video feature corresponding to the specified ranking and the positive sample data associated with the target video feature in the video feature set, the training evaluation method based on the recall model may further include steps S210 to S220, which are described in detail as follows:

[0087] Step S210, obtaining the recall rates calculated in different preset time intervals.

[0088] In this embodiment, the recall model is evaluated once every preset time interval to obtain the recall rate, so the recall rate calculated in different preset time intervals can be obtained. For example, the recall rate calculated in 3 different preset time intervals can be obtained, the recall rate calculated in 4 different preset time intervals can be obtained, and so on.

[0089] Step S220 , compare the values ​​of the multiple recall rates obtained, and select the recall model version corresponding to the recall rate with the largest value as the recall model with the best training effect to be applied to the information recommendation system.

[0090] After obtaining the recall rates calculated within different preset time intervals, the obtained recall rate values ​​are compared to determine the recall rate with the largest value, i.e., the maximum recall rate. Since the larger the recall rate, the better the effect of representing the recall model, the recall model corresponding to the maximum recall rate can be used as the recall model with the best training effect and applied to the information recommendation system.

[0091] In this embodiment, the recall rates calculated within different preset time intervals are obtained, the numerical values ​​of the multiple recall rates obtained are compared, and the recall model version corresponding to the recall rate with the largest numerical value is selected as the recall model with the best training effect and applied to the information recommendation system, thereby improving the iteration efficiency of the recall model.

[0092] In an exemplary embodiment, Figure 2 In step S130 of the illustrated embodiment, the process of respectively calculating the degree of match between the target account feature and each target video feature contained in the video feature set may also include: performing vector inner product operations on the target account feature and each target video feature, and using the obtained operation as the degree of match between the target account feature and the corresponding target video feature.

[0093] When extracting account features and video features, the recall model can map the account feature data and video feature data to the same space, thereby obtaining account features and video features in vector form. In order to determine the matching degree between the target account features and the target video features, in this embodiment, vector inner product operations can be performed on the target account features and each target video feature, and the obtained operation results are used as the matching degree between the target account features and the corresponding target video features.

[0094] In this embodiment, the matching degree between the target account feature and each target video feature is determined by a vector inner product operation, thereby improving the calculation speed and accuracy.

[0095] The following is a detailed description of a specific application scenario of the embodiment of the present application. Figure 5 As shown, online training or evaluation of the recall model can include the following processes:

[0096] Real-time messages: During the operation of the video platform, real-time messages will be generated. In order to perform online training on the information recommendation model, real-time messages will be obtained in this embodiment.

[0097] Real-time data processing: After obtaining real-time messages, the real-time messages will be processed to obtain the data of the service object from the real-time messages. The data of the service object includes but is not limited to the attribute information of the service object, the behavior information of the service object, etc. The attribute information of the service object is not limited to age, gender, etc. The behavior information of the service object includes but is not limited to clicks, views, likes, comments, forwarding and other behaviors. It should be noted that in this application, the attribute information, behavior information and other data related to the service object are involved. When the above embodiments of this application are applied to specific products or technologies, they are all obtained with the permission or consent of the service object, and the extraction, use and processing of the relevant data comply with national safety standards and national laws and regulations.

[0098] Extracting and splicing features: After the real-time message is processed to obtain the data of the service object, in this embodiment, features can also be extracted from the data and spliced.

[0099] Construction of positive and negative samples: After extracting and concatenating features, positive and negative samples can be constructed based on the obtained features. The construction method can be flexibly set according to actual needs. For example, positive samples can include videos that have been watched for a certain period of time by the service object, videos that have been forwarded by the service object, etc.

[0100] Offline Sample Center: The offline sample center can randomly select videos from the candidate pool as negative samples and input the negative samples.

[0101] Online training of recall model: After constructing positive and negative samples, the recall model can be trained online based on the positive and negative samples.

[0102] Account Tower DNN extracts account features: The recall model includes the Account Tower DNN module. During the online training of the recall model, the Account Tower DNN module extracts features from positive and negative samples to obtain account features in the form of vectors. The Account Tower DNN module can extract features from newly added positive and negative samples at regular intervals, and the time interval can be in minutes.

[0103] Video Tower DNN extracts video features: The recall model includes Video Tower DNN, which can extract features from the videos contained in the candidate pool to obtain video features in the form of vectors. The videos in the candidate pool are dynamically changing, and Video Tower DNN can extract features from the new version of the candidate pool at regular intervals, and the time interval can be in minutes.

[0104] Online index update: After obtaining the video features of the videos in the candidate pool, the video features can be imported into a search library, such as Faiss, which is an open source clustering and similarity search library developed by the Facebook AI team.

[0105] Online service: The online service searches for videos that match the account features from the search library.

[0106] Storing account features: The account tower DNN module extracts features from positive and negative samples. After obtaining the account features, in this embodiment, the obtained account features are stored in the account feature library, and the time when the account features are obtained is recorded.

[0107] Storing video features: Video Tower DNN can extract features from the videos contained in the candidate pool. After obtaining the video features, it stores the obtained video features in the account feature library and records the acquisition time of the video features corresponding to each version of the cache pool.

[0108] Sampling comparison: In the account feature library, the target account feature is sampled offline from the account feature corresponding to the time at every preset time interval, and the acquisition time of the target account feature is determined. The preset time interval and sampling amount can be set arbitrarily. For example, the preset time can be 1 hour, and the sampling amount can be 100,000. In other words, 100,000 target account features are sampled every 1 hour. Then, for each target account feature, the video feature corresponding to the candidate pool whose acquisition time is earlier than the acquisition time of the target account feature and closest to the acquisition time of the target account feature is searched in the video feature library. The video feature corresponding to the candidate pool is used as the video feature set of the target account feature. The matching degree between the target account feature and each target video feature contained in the corresponding video feature set is calculated respectively, and the top K target video features are selected based on the sorting of the matching degree values ​​from large to small. The value of K can be flexibly set according to actual needs. For example, the value of K can be determined by calculating the AA experiment fluctuation degree of the recall index through multiple random sampling.

[0109] Calculate the recall rate: determine the number of positive samples in the videos corresponding to the top K target video features; and determine the number of positive samples in the videos corresponding to the video feature set; obtain the ratio of the two to get the recall rate, and average the recall rate corresponding to each target account feature selected within the time interval to obtain the recall rate within the time interval.

[0110] Recall rate storage: After obtaining the recall rate of each time interval, store the recall rate of each time interval, the name of the recall model, the time interval, the number of target account features sampled in the time interval, and the training video samples. The storage can be in a table or a distributed file system.

[0111] Model comparison: Based on the stored recall rates, you can compare the recall models of different time intervals and the recall rates of different recall models, so as to select the recall model with the best effect to go online and improve iteration efficiency. In addition, you can also aggregate the recall rates of the recall models at different time intervals.

[0112] See also Figure 6 , Figure 6 FIG. 1 is a block diagram of a training evaluation device based on a recall model, shown in an exemplary embodiment of the present application. Figure 6 As shown, the device comprises:

[0113] A feature acquisition module 610 is configured to acquire account features and video features extracted based on training video samples during the online training of the recall model, and store the account features and video features in an account feature library and a video feature library, respectively;

[0114] The offline sampling module 620 is configured to obtain target account features from the account feature library by offline sampling, and search the video feature library for a video feature set related to the storage time according to the target account features;

[0115] A matching ranking module 630 is configured to respectively calculate the matching degree between the target account feature and each target video feature contained in the video feature set, and select a designated ranking based on the matching degree values ​​sorted from large to small;

[0116] The recall rate calculation module 640 is configured to calculate the recall rate based on the number of positive samples associated with the target video feature corresponding to the specified ranking and the positive sample data associated with the target video feature in the video feature set. The recall rate is used to evaluate the training effect of the recall model.

[0117] In another exemplary embodiment, the feature acquisition module 610 is also configured to record the acquisition time of the account features and the acquisition time of the video features in the account feature library and the video feature library respectively during the process of storing the account features and the video features in the account feature library and the video feature library respectively.

[0118] In another exemplary embodiment, the offline sampling module 620 includes:

[0119] A feature sampling unit, which samples target account features from the account feature library offline and determines the acquisition time corresponding to the target account features;

[0120] The feature search unit is configured to search the video feature library for a video feature whose acquisition time is earlier than the acquisition time of the target account feature and is closest to the acquisition time of the target account feature, and use the searched video feature as the target video feature in the video feature set.

[0121] In another exemplary embodiment, the feature sampling unit is further configured to periodically sample a specified number of target account features from the account feature library offline based on a preset time interval, wherein the preset time interval is greater than the frequency of extracting account features and video features during the online training process of the recall model.

[0122] In another exemplary embodiment, the apparatus further comprises:

[0123] A recall rate acquisition subunit, configured to acquire the recall rates calculated at different preset time intervals;

[0124] The recall rate comparison subunit is configured to compare the numerical values ​​of the multiple recall rates obtained, and select the recall model version corresponding to the recall rate with the largest numerical value as the recall model with the best training effect to be applied to the information recommendation system.

[0125] In another exemplary embodiment, the matching ranking module 630 is configured to perform vector inner product operations on the target account feature and each target video feature respectively, and use the obtained operation as the matching degree between the target account feature and the corresponding target video feature.

[0126] It should be noted that the recall model-based training evaluation device provided in the above embodiment and the recall model-based training evaluation method provided in the above embodiment belong to the same concept, and the specific manner in which each module and unit performs operations has been described in detail in the method embodiment and will not be repeated here.

[0127] An embodiment of the present application also provides an electronic device, including: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by one or more processors, the electronic device implements the methods provided in the above embodiments.

[0128] Figure 7 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application is shown.

[0129] It should be noted that Figure 7 The computer system 1600 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0130] like Figure 7 As shown, the computer system 1600 includes a central processing unit (CPU) 1601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1602 or the program loaded from the storage part 1608 to the random access memory (RAM) 1603, such as executing the method described in the above embodiment. In RAM 1603, various programs and data required for system operation are also stored. CPU 1601, ROM 1602 and RAM 1603 are connected to each other through bus 1604. Input / output (I / O) interface 1605 is also connected to bus 1604.

[0131] The following components are connected to the I / O interface 1605: an input section 1606 including a keyboard, a mouse, etc.; an output section 1607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1608 including a hard disk, etc.; and a communication section 1609 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1609 performs communication processing via a network such as the Internet. A drive 1610 is also connected to the I / O interface 1605 as needed. A removable medium 1611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1610 as needed so that a computer program read therefrom is installed into the storage section 1608 as needed.

[0132] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication section 1609, and / or installed from a removable medium 1611. When the computer program is executed by a central processing unit (CPU) 1601, various functions defined in the system of the present application are executed.

[0133] It should be noted that the computer-readable medium shown in the embodiment of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as a part of a carrier wave, wherein a computer-readable computer program is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0134] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the system, method and computer program product according to various embodiments of the present application. Wherein, each box in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0135] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. The names of these units do not, in some cases, constitute limitations on the units themselves.

[0136] Another aspect of the present application further provides a computer-readable storage medium on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor of an electronic device, the electronic device implements the method described above. The computer-readable storage medium may be included in the electronic device described in the above embodiment, or may exist independently without being assembled into the electronic device.

[0137] Another aspect of the present application further provides a computer program product or a computer program, which includes a computer instruction, and when the computer instruction is executed by a processor, the method provided in each of the above embodiments is implemented. The computer instruction can be stored in a computer-readable storage medium; the processor of the electronic device can read the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the electronic device performs the method provided in each of the above embodiments.

[0138] The above content is only a preferred exemplary embodiment of the present application and is not intended to limit the implementation scheme of the present application. A person skilled in the art can easily make corresponding changes or modifications based on the main concept and spirit of the present application. Therefore, the scope of protection of the present application shall be based on the scope of protection required by the claims.

Claims

1. A training evaluation method based on a recall model, characterized in that: include: Acquire the account features and video features extracted based on the training video samples during the online training of the recall model, and store the account features and the video features in an account feature library and a video feature library respectively; Offline sampling is performed from the account feature library to obtain target account features, and a video feature set related to the storage time is searched from the video feature library for the target account features; Calculating the matching degree between the target account feature and each target video feature contained in the video feature set respectively, and selecting a designated ranking based on the sorting of the matching degree values ​​from large to small; A recall rate is calculated according to the number of positive samples associated with the target video feature corresponding to the specified ranking and the positive sample data associated with the target video feature in the video feature set, and the recall rate is used to evaluate the training effect of the recall model.

2. The method according to claim 1, characterized in that The method further comprises: In the process of storing the account features and the video features in the account feature library and the video feature library respectively, the acquisition time of the account features and the acquisition time of the video features are also recorded in the account feature library and the video feature library respectively.

3. The method according to claim 2, characterized in that The offline sampling from the account feature library to obtain the target account feature, and searching the video feature set related to the storage time from the video feature library for the target account feature, includes: Offline sampling of target account features from the account feature library, and determining acquisition time corresponding to the target account features; The video feature library is searched for a video feature whose acquisition time is earlier than the acquisition time of the target account feature and is closest to the acquisition time of the target account feature, and the searched video feature is used as the target video feature in the video feature set.

4. The method according to claim 3, characterized in that The offline sampling of target account features from the account feature library includes: A specified number of target account features are periodically sampled offline from the account feature library based on a preset time interval, wherein the preset time interval is greater than a frequency of extracting account features and video features during online training of the recall model.

5. The method according to claim 4, characterized in that After calculating the recall rate according to the number of positive samples associated with the target video feature corresponding to the specified ranking and the positive sample data associated with the target video feature in the video feature set, the method further includes: Get the recall calculated at different preset time intervals; The obtained multiple recall rates are compared in numerical value, and the recall model version corresponding to the recall rate with the largest value is selected as the recall model with the best training effect and applied to the information recommendation system.

6. The method according to any one of claims 1 to 4, characterized in that The respectively calculating the matching degree between the target account feature and each target video feature contained in the video feature set includes: A vector inner product operation is performed on the target account feature and each target video feature respectively, and the obtained operation is used as the matching degree between the target account feature and the corresponding target video feature.

7. The method according to any one of claims 1 to 4, characterized in that The type of the account feature library includes a real-time table, and the type of the video feature library includes a distributed file system.

8. A training evaluation device based on a recall model, characterized in that: include: A feature acquisition module, configured to acquire account features and video features extracted based on training video samples during the online training of the recall model, and store the account features and video features in an account feature library and a video feature library, respectively; An offline sampling module, configured to obtain target account features from the account feature library by offline sampling, and search the video feature set related to the storage time from the video feature library for the target account features; A matching ranking module, configured to respectively calculate the matching degree between the target account feature and each target video feature contained in the video feature set, and select a designated ranking based on the order of matching degree values ​​from large to small; The recall rate calculation module is configured to calculate the recall rate based on the number of positive samples associated with the target video feature corresponding to the specified ranking and the positive sample data associated with the target video feature in the video feature set, and the recall rate is used to evaluate the training effect of the recall model.

9. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the electronic device to implement the recall model-based training and evaluation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Computer-readable instructions are stored thereon, and when the computer-readable instructions are executed by a processor of a computer, the computer is caused to execute the recall model-based training and evaluation method described in any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the recall model-based training and evaluation method described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Unified text analytics annotator development life cycle combining rule-based and machine learning based techniques

    US20180246867A1

  • Real-time drift detection in machine learning systems and applications

    US20200082296A1